Jacky Liang placed 11 large language models in a 2D battle royale simulation across 30 matches, spending $482 on inference to test real-world competitive performance. Grok 4.1 Fast emerged victorious with a 43% win rate at just $0.97 per win, while Claude prioritized cooperation and Grok adopted aggressive tactics. The experiment demonstrates that traditional benchmarks don’t predict real-world performance and that model selection should depend on task requirements rather than benchmark rankings alone.