Authors: Ali Mahmoudzadeh, Salma Elshafey, Shuo Qiu, José Santos, Ilya Matiach, Vivek Bhadauria, Morteza Ziyadi, April Kwong LLM-as-judge evaluations are a critical part of any agent lifecycle. Single-turn metrics such as relevance, faithfulness, tone… can be straight forward to prompt and analyze, but a number per response isn’t the thing your users experience. They experience a session : a multi-turn conversation in which the agent asks clarifying questions, calls tools, retries when something fails, and hopefully stitches the whole thing into a usable outcome.