OpenAI’s FrontierScience benchmark reveals the gap between structured problems and real research work. GPT-5.2 scores 77% on Olympiad tasks but only 25% on open-ended research—key insights for AI support teams.
Need help?
Contact usOpenAI’s FrontierScience benchmark reveals the gap between structured problems and real research work. GPT-5.2 scores 77% on Olympiad tasks but only 25% on open-ended research—key insights for AI support teams.