Papers for

healthcare software engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Self-training causes performance decline in biomedical question answering AI

Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study

Abstract: Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.

Mon 28 SeptComputation and LanguageMachine Learning
The gist
When AI models are trained repeatedly using their own answers, mistakes and biases can pile up over time. The authors tested this effect on models answering biomedical questions and found that their accuracy and answer quality get worse with each new generation of training. Larger models showed a bigger drop in performance than smaller ones. This work helps understand risks when AI learns from its own predictions in specialized areas like medicine.
Open → 2609.34257v1

Large language models improve traditional chinese medicine prescriptions with safer reasoning

Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation

Abstract: Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao paradigm (SR Gap), treat each encounter in isolation without follow-up adjustment via sui zheng jia jian (LA Gap), and fail to enforce absolute contraindication rules such as Shi Ba Fan (SC Gap). We propose a progressive four-stage framework (SFT $\to$ PG-CoT $\to$ Dynamic $\to$ K-RL) that addresses each gap: PG-CoT constrains CoT distillation under the li-fa-fang-yao paradigm to produce auditable diagnostic chains, Dynamic SFT models patient trajectories with explicit transition reasoning, and K-RL encodes deterministic pharmacological rules as rule-based DPO preference signals. Across 12 fine-tuned models and 6 zero-shot baselines, our framework substantially improves prescription quality over zero-shot baselines---with a 7B model (Mistral-7B) surpassing zero-shot GPT-5 on all three TCM evaluation metrics.

Tue 22 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models can create Traditional Chinese Medicine (TCM) prescriptions, but they often miss important reasoning steps and safety rules. The authors found three main problems: no clear reasoning process, ignoring changes in patient follow-ups, and failing to avoid harmful ingredient combinations. They built a multi-stage method that teaches models to explain their reasoning, consider patient history, and follow strict safety rules. This approach made the models produce better and safer TCM prescriptions than previous versions.
Open → 2609.25755v1