A team routing queries across a coding specialist, a logic specialist, and a generalist model assumes each will cover the others’ blind spots. A new study evaluating 67 frontier models from 21 providers shows that assumption is mathematically flawed — and the flaw has a name: the co-failure ceiling. The assumption works like this: as long as two models don’t usually fail on the exact same prompts, combining them is supposed to create a safety net against failures.