Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models
2026-08-31 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors look at how teaching formats affect updating medical knowledge in large language models, which can struggle with old but believable information. They create a new test called SEER-Bench using recent cancer data and try four teaching styles to update models. They find that the EMQ format works best for keeping knowledge accurate and stable. Their analysis shows EMQ helps the model learn clear medical distinctions without drastically changing its original understanding. Overall, the way information is presented matters a lot when updating medical knowledge in AI.
large language modelsmedical knowledge updatingsupervision formatSEER-Benchoncology stagingEMQNCCN guidelinesfine-tuningmodel retentionclinical contrast
Authors
Yangmin Huang, Shu Quan, He Geng, Xin Ye, Qianyun Du, Zhiyang He, Jiaxue Hu, Xiaodong Tao
Abstract
Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget setting. We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats: EMQ, MSQ, FITB, and SAQ. Across SEER-Bench and HealthBench Professional, EMQ gives the most stable external transfer and retention among same-budget SFT variants. With EMQ supervision, the updated 4B model produces competitive results on temporally anchored oncology staging, reaching 64.8% answer accuracy and 59.6% rationale accuracy on SEER-Bench. Diagnostic analyses suggest that EMQ exposes denser clinical contrast signals while preserving discriminative representations with smaller movement from the base model. These results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision.