Yesterday, Hugging Face came out saying they’d detected an AI autonomous-agent-powered cyberattack and that they had to use open-source models to actually investigate and remediate it. Later we heard from OpenAI that their agent was responsible; it happened during an ExploitGym eval, and the agent just drifted off the goal. It escaped the sandbox, reached OpenAI Research Environment, got access to internet, and hacked Hugging Face production environment trying to find the solution for the benchm