Story perspectives
AI Models Score Just 25% on Complex White-Collar Tasks
1/23/2026
1 of 1
Story summary
- Mercor's APEX-Agents benchmark shows AI models score about 25% on complex white-collar tasks in consulting and law.
- The benchmark targets multi-domain reasoning essential for knowledge work and reflects professional tasks.
- Gemini 3 Flash and GPT-5.2 led with 24% and 23% accuracy, while no model can replace investment bankers.
- Mercor CEO Brendan Foody says results contrast with OpenAI's GDPval, which tests general knowledge, and Foody anticipates AI improvements.
