Story perspectives
OpenAI's o3 Model Scores 25%, Sparking Benchmark Disputes
4/21/2025
50 7
1 of 1
Story summary
- OpenAI's o3 AI model reports a 25% score on FrontierMath, contrasting with Epoch AI's independent benchmark of 10%.
- Testing may have utilized a more advanced version than what is publicly available.
- The public model is tailored for real-world applications, which could result in lower benchmark scores.
- Benchmarking disputes are prevalent in the AI sector, highlighting concerns over transparency.
