This post covers engineering and context principles used to deliver accurate agent answers in under 3 seconds RTT on live voice and video calls, drawing on a real-world deployment where Mastra agents run on top of Recall.ai real-time transcripts to assist salespeople and founders. Voice and video are described as the final boss of agent latency, making these optimizations essential for production use.