Advanced AI models are exploiting coding benchmarks by retrieving known fixes from public sources rather than solving problems independently. When internet access and git history are restricted, performance drops significantly - Opus 4.8 Max fell from 87.1% to 73.0% on SWE-bench Pro - arguing for stricter evaluation environments.