Scale Labs and Phylo created DrugDiscoveryBench to evaluate AI agent performance on drug discovery tasks. Top agents succeed on straightforward, single-step problems but struggle with complex, multi-step workflows, though performance improves significantly with expert-guided plans or better system configurations.