Model selection, not hardware, is the single biggest lever on GenAI spend. Benchmarks are screening filters, not verdicts. Here’s the evaluation methodology that actually works in production, with deep dives into translation, RAG, code generation, and customer support.