ForgeCode hit 78.4% SOTA on TermBench 2.0 with gemini-3.1-pro-preview. This is the technical account of how we got there: seven failure modes, their fixes, and why the benchmark work generalized across models rather than overfitting to one run.
Benchmarks Don't Matter — Until They Do (Part 1)
calendar_today
March 3, 2026
domain
tailcall