Story perspectives
AI Fails Real-World Medical Tests, Study Reveals Shortcomings
8/24/2025
42 9
1 of 1
Story summary
- A study in JAMA Network Open shows AI systems excel in tests but falter in real-world medical reasoning.
- Performance significantly declines when familiar patterns change, highlighting a reliance on pattern recognition.
- Six AI models, including GPT-4o, exhibited notable accuracy drops, with Llama 3.3-70B showing nearly 40% more errors.
- The study calls for better evaluation tools and AI models for complex medical situations.
