Automated benchmark improves evaluation of clinical ai in health records
A Living Benchmark for Information Retrieval from Electronic Health Records
Artificial Intelligence
Summary
Finding important information in electronic health records (EHRs) can be hard and time-consuming for doctors. The authors created a tool that automatically makes question-and-answer pairs based on patient notes to test how well AI systems can find needed info. Nineteen clinicians checked this tool and helped build a benchmark called BRIE that can keep updating itself as new data comes in. Testing showed current AI models often miss key details, especially when answers require combining information from different visits. This approach helps keep AI clinical tools tested and safe as medical records and technology change.
What this means in practice
- •For ehr system developers: Use the BRIE benchmark to continuously test and improve AI assistants for retrieving accurate patient information from health records.$Commercial implications: Enables creation and marketing of safer, more reliable AI clinical assistant products integrated in EHR systems.
- •For healthcare quality assurance teams: Deploy scalable benchmarks like BRIE to regularly validate AI tools that support clinician decision-making with up-to-date patient information.
Authors
Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
Abstract
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.