Story perspectives
AI Excels in Medical Exams, Struggles with Patient Interactions
1/3/2025
29 7
1 of 1
Story summary
- AI models are proving their prowess in medical exams, yet they falter in patient interactions, leading to a dramatic drop in diagnostic accuracy. For instance, GPT-4's accuracy plummeted from 82% in structured cases to just 26% during simulated conversations. The CRAFT-MD benchmark underscores the critical need for real-life scenarios to assess AI's clinical reasoning and its ability to gather comprehensive medical histories.
