A benchmark evaluating AI agents on long-horizon, economically valuable, real-world tasks across 55 professional fields, drawing on 300+ experts. The hardest tier shows frontier models averaging below 1% full pass rate.
Need help?
Contact usA benchmark evaluating AI agents on long-horizon, economically valuable, real-world tasks across 55 professional fields, drawing on 300+ experts. The hardest tier shows frontier models averaging below 1% full pass rate.