Story perspectives
xAI Faces Scrutiny Over Misleading AI Benchmark Claims
2/23/2025
32 7
1 of 1
Story summary
- An OpenAI employee has accused xAI of presenting misleading benchmark results for Grok 3.
- Igor Babushkin from xAI defended the company's accuracy in reporting.
- The legitimacy of AIME 2025 as an AI benchmark is under scrutiny.
- xAI's graph excluded OpenAI's o3-mini-high score, potentially skewing performance representation.
- AI benchmarks frequently do not disclose models' actual limitations and associated costs.
