Through collaboration with Anyscale, Google introduced three architectural optimizations to Ray Serve LLM on Kubernetes: HAProxy integration for request routing, direct token streaming, and an updated Ray executor backend for vLLM, delivering up to 5x higher throughput and 8x lower latency. The improvements were benchmarked on GKE clusters using NVIDIA HGX B200 systems running Gemma 4 E2B, with developers encouraged to try Ray 2.56 and later.
Scaling Ray Serve LLM on GKE: Performance without losing the developer experience
calendar_today
June 18, 2026
domain
google-pub-sub