Story perspectives
AI Safety Tests Exposed: 440 Flaws Found, Standards Needed
11/5/2025
38 3
1 of 1
Story summary
- Experts identified flaws in more than 440 AI safety and performance tests.
- A study by the British government's AI Security Institute and the University of Oxford and the University of California, Berkeley found that nearly all benchmarks have weaknesses.
- Google withdrew its AI model Gemma after it generated false allegations about a U.S. senator.
- The study calls for shared AI testing standards, noting only 16% of benchmarks used uncertainty estimates.
