This guide explains how enterprise teams can move beyond simplistic LLM metrics like tokens-per-second to evaluate inference through multidimensional performance trade-offs. It argues that the two metrics vendors highlight on landing pages, tokens per second and cost per million tokens, fail to capture real production behavior.