At the PyTorch Conference 2025, we demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters using Primus-Turbo, an AMD optimization library for training frameworks such as TorchTitan. We have since upstreamed those AMD optimizations so TorchTitan supports AMD Instinct(™) GPUs directly, with competitive FP8 performance out of the box. All contributions mentioned have been merged into upstream pytorch/AO and pytorch/TorchTitan.
FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance Improvements
calendar_today
August 13, 2026
person
AMD: Rishi Sinha, Yuankai Chen, Liz Li, Shekhar Pandey, Wen Chen, Xiaobo Chen, Yao Fu, Zhenyu Gu, Andy Luo, Peng Sun META: Matthias Reso, Hamid Shojanazeri, TorchAO team, TorchTitan team
domain
pytorch