By Alex Ginglen Imagine you’re in charge of building a generative AI agent chatbot. You’ve done countless user testing sessions before deploying to production. You’ve even written custom regression tests using LLM as a Judge and BERTScore that runs in your CI/CD pipeline.