TL;DR Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments show that Helion provides a productive PyTorch-native…
Portable vLLM Model Inference Kernels in Helion
calendar_today
June 10, 2026
person
Sean Chen (Red Hat) and Yanan Cao (PyTorch, Meta Platforms)
domain
pytorch