In the last two months, there have been announcements for several landmark models: GPT-4.1 on April 14th, Llama 4 on April 5th, Gemini 2.5 on March 25th, and Claude 3.7 at the end of February. Undoubtedly, there’s a frenetic amount of work going into training the next generation of foundation models, and everything is changing fast. This continuous change is a great reminder that evaluating LLMs…