Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed training across hundreds or thousands of GPUs. At this scale, hardware failures are not exceptional events—they are expected….
Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm
calendar_today
July 6, 2026
person
AMD: Chaojun Hou, Liz Li, Zachary Streeter, Xinyu Kang, Lei Zhang, Yuankai Chen, Yao Fu, Wen Chen, Zhenyu Gu, Andy Luo │ Meta: Matthias Reso, Hamid Shojanazeri, Monarch team
domain
pytorch