If you run GPU workloads on Kubernetes — vLLM, Triton, training jobs, or the newer agentic inference stacks — you’ve probably hit a familiar problem: the default autoscaling path still reasons about CPU and memory, while…
GPU autoscaling on Kubernetes with KEDA: Building an external scaler
calendar_today
May 27, 2026
person
Pavan Madduri (Senior Cloud Platform Engineer @ Grainger | CNCF Golden Kubestronaut)
domain
cncf