1 of 1
Story summary
- A global consortium of nearly 1,000 experts developed Humanity’s Last Exam (HLE), a 2,500-question assessment to evaluate advanced AI capabilities.
- Unlike traditional benchmarks, HLE focuses on specialized knowledge across fields such as ancient languages and advanced mathematics.
- Early results show GPT-4o and Claude 3.5 scoring 2.7% and 4.1%, and Dr. Tung Nguyen of Texas A&M says HLE reveals AI limitations and the value of human knowledge.
