NVIDIA unveiled Nemotron 3 Ultra, a 550-billion-parameter mixture-of-experts model designed to enhance long-running agent workflows. The model achieves 5x higher throughput compared to other open models in its class while reducing task completion costs by up to 30% through efficient token usage. The release includes hybrid Mamba-Transformer layers, NVFP4 quantization for cross-GPU deployment, and Multi-Teacher On-Policy Distillation training leveraging over ten specialized teacher models.