What closed and open models actually do when you tell them to hack a website Summary The cybersecurity capability of a model is currently measured by a solve rate, a percentage, and that number tells you almost nothing worth knowing. It doesn’t tell you how the model behaved. It doesn’t tell you what it is actually capable of, or where it is lacking, or how far it still is from doing the job end to end.
Watching Agents Work: A Behavioral Audit of Offensive-Security LLM Runs
calendar_today
August 3, 2026
domain
nuclei