Self-training causes performance decline in biomedical question answering AI
Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
Computation and LanguageMachine Learning
Summary
When AI models are trained repeatedly using their own answers, mistakes and biases can pile up over time. The authors tested this effect on models answering biomedical questions and found that their accuracy and answer quality get worse with each new generation of training. Larger models showed a bigger drop in performance than smaller ones. This work helps understand risks when AI learns from its own predictions in specialized areas like medicine.
What this means in practice
- •For biomedical ai developers: Identify risks of recursive self-training causing performance drops and biases when developing biomedical QA models.
- •For healthcare software engineers: Avoid recursive synthetic data training cycles that degrade accuracy in biomedical question answering applications.
Tested on one dataset.
Authors
Bibek Bhandari, Kshitij Lingthep
Abstract
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.