Medical studies fall behind fast changing clinical ai models
The widening evaluation gap in medical large language model research 2023 to 2026
Computation and Language
Summary
Medical AI models improve very quickly, but studies testing them take much longer to complete. The authors found that evaluations of these models are getting further and further behind the latest versions. Trials that use strict methods tend to test older models than other study types. This means there’s a trade-off between testing the newest AI and using rigorous study designs. The gap comes mostly from which AI models researchers choose to evaluate, not how long studies take.
What this means in practice
- •For clinical trial designers: Plan trials considering the model age gap to balance rigour with evaluating current AI technologies.
- •For healthcare ai product managers: Adapt product evaluation timelines to account for delays in clinical study publications on AI models.
Authors
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
Abstract
Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.