Story perspectives
OpenAI Launches FrontierScience Benchmark for AI Research Evaluation
12/16/2025
1 of 1
Story summary
- OpenAI introduced FrontierScience, a benchmark to evaluate AI models' capabilities in scientific research.
- The FrontierScience benchmark includes two tiers in physics, chemistry, and biology.
- GPT-5.2 achieved 77.1% on the Olympiad tier and 25.3% on the Research tier.
- Researchers say the benchmark advances the field but does not assess experimental execution or image analysis.
- AI advances, particularly in reinforcement learning, are noted, yet reliable evaluation of specialized knowledge remains challenging.
