A few weeks ago, an AI cyber evaluation produced an unexpectedly efficient strategy for solving a benchmark: the agents went looking for the answers. According to OpenAI’s preliminary disclosure, models being tested for advanced cyber capabilities found ways to obtain secret information that could help them complete a benchmark. They chained vulnerabilities, stolen credentials, internet access, and inferences about where benchmark material might be hosted.