A Google and Anyscale partnership delivers up to 5x higher throughput and 8x lower latency in Ray Serve through three architectural optimizations: HAProxy integration, direct token streaming, and a v2 Ray executor backend for vLLM. Benchmarks demonstrate improved scaling on GKE clusters with NVIDIA HGX B200 systems.