Today weβre open-sourcing hyper-π-bench, a new long horizon agent evaluation that measures how well models can not only act as an agent, but construct one.
Need help?
Contact usToday weβre open-sourcing hyper-π-bench, a new long horizon agent evaluation that measures how well models can not only act as an agent, but construct one.