Drooid Logo
Back to today’s briefing

Story perspectives

AI Models Score Just 25% on Complex White-Collar Tasks

1/23/2026

31 7 Full Breakdown

1 of 1

Story summary
  • Mercor's APEX-Agents benchmark shows AI models score about 25% on complex white-collar tasks in consulting and law.
  • The benchmark targets multi-domain reasoning essential for knowledge work and reflects professional tasks.
  • Gemini 3 Flash and GPT-5.2 led with 24% and 23% accuracy, while no model can replace investment bankers.
  • Mercor CEO Brendan Foody says results contrast with OpenAI's GDPval, which tests general knowledge, and Foody anticipates AI improvements.