Baseten describes serving GLM-5.2 at 280+ tokens per second on NVIDIA Blackwell hardware using KV-aware routing, PD disaggregation, Multi-Token Prediction, NVFP4, and other optimizations.
Need help?
Contact usBaseten describes serving GLM-5.2 at 280+ tokens per second on NVIDIA Blackwell hardware using KV-aware routing, PD disaggregation, Multi-Token Prediction, NVFP4, and other optimizations.