Authors: Caroline Yeh, Ali Mahmoudzadeh, Morteza Ziyadi If you’ve shipped an agent to production, or you’re close, you’ve already made a quiet bet: that the evals you ran during development reflect what happens when real users interact with the real system. That bet is harder to win than it looks. Not because evaluation is unsolved, but because agents break in ways single-turn metrics can’t see.