New benchmark evaluates AI for everyday patient care
Mass General Brigham has developed a new benchmarking approach to evaluate the real-world performance of AI models in clinical settings. Unlike standard tests, this method assesses AI tools on routine patient care tasks using actual hospital data, aiming to ensure reliability and safety before widespread deployment.
Background
- Mass General Brigham is a major US hospital system that runs Harvard-affiliated hospitals and regularly publishes health-tech research.
- Large Language Models (LLMs) like GPT-4 are AI systems trained on text; in medicine, they are being tested for tasks like triaging messages, answering patient questions, or summarizing records.
- This press release announces a new benchmark (meaning a standardized test to measure performance) called "CareBench," designed specifically for the routine, non-specialist tasks that primary care doctors handle daily — as opposed to existing medical benchmarks that focus on narrow specialties like radiology or pathology.
- The researchers tested 7 leading LLMs and found that even the best ones (Claude 3.5 Sonnet, GPT-4) made frequent mistakes on everyday primary care scenarios, such as determining appropriate screening intervals or follow-up plans, raising questions about how AI can safely be used in general clinical settings.