This research examines optimization strategies for search agents across the retrieval tool and the planning mechanism, finding that the reranker matters most for tool configuration. Training the planner through on-policy distillation combined with a novel efficiency reward (CLP) enables a faster-configured system to match stronger baselines at half the latency while reaching 60.7% accuracy on their primary benchmark.
Making Search Agents Faster and Smarter
calendar_today
March 4, 2026
person
Abdallah Bashir, Bo Han, Fanhai Lu, Sheshansh Agrawal
domain
contextual-ai