Evaluation methods guide healthcare large language model testing
A primer on evaluation methods for large language models in healthcare
Computation and LanguageArtificial Intelligence
Summary
Large language models are increasingly used in medicine, but making sure they work well is tricky because they give complex and changing answers. The authors review important ways to test these models, including study design, statistical methods, and checking their abilities through quizzes and simulated conversations. They also discuss how to judge the clinical accuracy of the models’ free text answers, such as having humans or the models themselves evaluate outputs and running clinical trials. Their goal is to help people design better tests to make sure these AI tools help rather than harm patients.
What this means in practice
- •For healthcare ai developers: Design rigorous tests to evaluate large language models for medical applications using a variety of benchmarks and clinical validation approaches.
- •For clinical trial designers: Incorporate appropriate evaluation methods for free-text AI outputs to ensure safety and accuracy in medical trials involving language models.
A survey. It maps existing work.
Authors
Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries, Alyssa Pradhan, Timothy J Keyes, Bernardo C Bizzo, Keith J Dreyer, Sarah F Mercaldo, James M Hillis
Abstract
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.