Could an AI benchmark evaluate Machiavellian play? Coupbench tries: incomplete-information social games instead of static puzzles. Its preliminary chart shows Gemini 3.8 Flash at 76.6 vs Muse Spark 1.3 at 73.0—but the creator says it’s not statistically significant yet.