AI agents generate 10x to 100x more tokens than chatbots. Without optimization, inference costs dominate your cloud bill. This guide covers the four techniques that cut agent spend by 60 to 80 percent: model routing, prompt caching, context management, and budget caps.