A new benchmark finds that medical AI models perform well on standardized exams but struggle with real patient care, highlighting a gap between theoretical knowledge and practical clinical application.
Background
- The article contrasts how AI (like GPT-4) performs on medical licensing exams (USMLE) — where it scores near-perfect — versus a new benchmark called CRAFT-MD that tests real clinical decision-making (e.g., ordering labs, adjusting meds, knowing when to refer). AI's performance drops sharply on the latter.
- CRAFT-MD is introduced as a more realistic evaluation: it simulates patient visits with unfolding information over time, not just static questions. It was developed by researchers at Stanford, UC Berkeley, and others to reveal where models still fail: contextual reasoning, multi-step planning, and handling ambiguity.
- The gap matters because hospitals and clinics are already piloting AI for triage, clinical notes, and decision support. Over-reliance on exam scores could lead to deployment of tools that appear competent but miss subtleties in actual care.
- Also notable: the same pattern — high test scores, poor real-world transfer — has been observed in AI for law, finance, and software engineering, making this a broader AI evaluation problem.