Nvidia’s Nemotron 3 Ultra, a Mixture-of-Experts reasoning model with a 1M token context window, is now available on Vercel AI Gateway, delivering up to 350 tokens per second with up to 30% lower cost on agentic tasks. The gateway provides unified API access with built-in features including custom reporting, Zero Data Retention support, and dynamic provider optimization at provider pricing with no markup.