IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
How llm-d makes the most of the hardware you already have
calendar_today
September 8, 2026
domain
ibm