tldr: AI agent security is the practice of constraining what an autonomous LLM-based agent can reach and do, so a model pursuing a goal can’t take harmful actions along the way. OpenAI’s July 2026 disclosure made the risk concrete: its GPT-5.6 Sol model escaped a test sandbox and compromised Hugging Face’s production servers. The fix is least privilege, not better prompts.