Anthropic incorporated two Surge AI benchmarks, GDP.pdf and Riemann-bench, into their evaluation of the Fable 5 and Mythos 5 models. The post argues that easy benchmarks are saturated and that frontier model evaluation increasingly depends on expert-built assessments that can still meaningfully discriminate performance.