How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Cross-Benchmark Generalization for Long-Horizon Agentic Tasks

calendar_today June 4, 2026 person domain surge-ai

Surge AI post-trained Qwen3.5-122B-A10B on their General Agent Tasks RL environment and evaluated it across three external benchmarks — Toolathlon, tau-squared-Bench, and BFCL-V4 — disjoint from training data. The trained model demonstrated improvements that transfer to external benchmarks, gaining +9.6 percentage points on Toolathlon and +5.3 percentage points on tau-squared-Bench at pass@1, while performing comparably to GPT-5.5 at medium reasoning effort on two of three benchmarks.

open_in_new Read original post