We ran the same coding tasks with and without prebundled tooling, across multiple models and languages. Here’s what changed. Eval-driven development IDE-native search reduced latency, cost, and budget overruns. The comparison below uses paired task-level deltas. Aggregate medians and totals are shown for orientation. Budget overruns are tasks that exceeded the USD 0.50 per-task cap. 8.33% Median latency reduced 83.11s → 79.03s 16.44% P95 latency reduced 268.71s → 213.17s 5.60% Total cost reduced USD 44.17 → USD 41.67 33.28% Budget overruns reduced 6.67% → 4.