The AI code evaluation framework behind our open vs. frontier model test: rubric-based scoring, a blinded LLM judge, and validation on real SWE-bench data.
AI model routing: How we score code without running tests
calendar_today
July 15, 2026
domain
faros