How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Are Claude’s Models Actually Getting Better? I Instrumented Claude Code to Find Out

calendar_today June 25, 2026 person hello@signoz.io (SigNoz Inc) domain signoz

Benchmark scores tell you whether a model solved a task, not what it cost to get there. I instrumented Claude Code with OpenTelemetry and SigNoz to compare Claude Sonnet 4.6, Opus 4.7, and Opus 4.8 across accuracy, cost per solved task, tokens, cache utilization, error rate, and active time, on a fixed Terminal-Bench task suite.

open_in_new Read original post