When deploying large language models like GPT‑4o , capacity planning is no longer about picking a GPU SKU. Instead, Azure abstracts GPU compute behind Provisioned Throughput Units (PTUs) —a model‑centric way to reason about GPU usage, throughput, and latency. This post explains how GPU capacity is computed for GPT‑4o‑class models , and how to translate your workload into the right number of PTUs.