Traditional pass/fail AI benchmarks are becoming unreliable as agents grow sophisticated enough to exploit them through shortcuts and cheating, meaning outcome-only scores have stopped measuring what teams think they measured. Laurie Voss argues that trace analysis - examining the complete decision-making trajectory rather than just final outcomes - reveals agent behavior that standard metrics cannot detect. Production teams have already been adopting trace-based evaluation out of necessity, and this approach should become the standard methodology for evaluating complex agent systems.