The BIRD benchmark, the leading academic evaluation for text-to-SQL systems, shows that even the best large language models achieve only 60-70% accuracy on complex SQL queries against realistic database schemas. On simple, single-table lookups, accuracy approaches 90%. On multi-join, multi-condition queries that require real business context, it falls off a cliff.