Hybrid architectures that interleave Mamba-style SSM layers with standard full-attention (FA) layers — such as NVIDIA Nemotron-H — are gaining traction as a way to combine the linear-time…
Disaggregated Serving for Hybrid SSM Models in vLLM
calendar_today
April 21, 2026
person
Nicolò Lucchesi, Zhanqiu Hu (Red Hat), and the vLLM team
domain
vllm