Researchers show that serving AI models with llm-d can boost inference speeds by up to 5 times and double throughput — all while using heterogeneous GPUs.
Running AI on mixed hardware for speed and affordability
calendar_today
June 23, 2026
domain
ibm