LLMs have gotten ‘smarter’ at an astronomical rate, driven largely by one metric: context window size. This is how much data an LLM can effectively ‘think’ about at once. What was once a ceiling of 16,385 tokens for models like GPT-3.5 is now routinely measured in the hundreds of thousands—with some models, like those from Gemini, offering context windows of 1 million tokens or more.This context window revolution has profound implications for building sophisticated voice agents.