For AppSec work, the harness around a model matters more than which model is chosen: GPT-5.5 reached 71.4% on the UK AI Security Institute’s hardest evaluation category while Mythos reached 68.6%, and UCSB research found the best harness passed four times as many tests as the worst harness using the same Claude Opus 4.6 model. Mythos produced 3 false positives and 1 non-security bug out of 5 “confirmed vulnerabilities” when tested against cURL, while eight GPT-5.4-mini agents can match one GPT-5.5 agent at similar cost, illustrating that orchestration design outweighs model selection for vulnerability discovery.