Story perspectives
AI’s Last Exam: Top Models Only 3.3% Accurate
1/26/2025
42 9
1 of 2
Story summary
- "Humanity's Last Exam," crafted by the Center for AI Safety and Scale AI, puts AI models to the test with 3,000 intricate questions spanning multiple disciplines. Shockingly, top performers like GPT-4o achieved a mere 3.3% accuracy. This benchmark not only highlights the reasoning and ethical shortcomings of AI but also underscores the urgent need for advancements in technology.
1 / 2
