Biomedical language models often make up fake research references

Biomedical Reference Generation Remains Unreliable across 26 Large Language Models

Computation and Language

Summary

Many computer programs that write about medical research sometimes invent fake papers when asked for references. A study tested 26 such language models and found that more than half of their references were not real. Even the best models only got all the details right about half the time. This means people using these tools to help write medical papers need to double-check the references carefully.

What this means in practice

  • For biomedical writers: Verify references suggested by language models before including them in medical documents to avoid citing nonexistent papers.
  • For medical publishing teams: Develop workflows to check bibliographic details of references generated by AI tools to maintain accuracy in published articles.

Authors

Maxim Topaz, Zhihong Zhang, Nir Roguin, Pallavi Gupta, Zichao Li, Laura-Maria Peltonen

Abstract

Background. Large language models are increasingly used to help write biomedical text but may fabricate references to nonexistent work. How often large language models do so is not well characterized. Methods. We prompted 26 language models from eight developers (2023 to 2026) to supply a missing reference for each of 69 biomedical passages across ten domains. References were classified as verifiable (real paper with a resolving identifier), partial matches (real paper without a resolving identifier), fabricated (no matching indexed paper), or declined (the model refused to supply a reference). A reference was considered correct in every evaluated bibliographic field only when it was verifiable and its journal, year, and listed authors matched those of the cited paper. Results. Fabrication ranged from 10.2% (Claude Opus 4.8, which declined 52.1% of prompts) to 98.4% (Ministral 3B, which produced no verifiable reference). Claude Opus 4.6 and Claude Sonnet 4.5 produced similar proportions of verifiable references (77.6% and 76.6%) but named authors correctly in 78.7% and 28.7% of author-evaluable verifiable references, respectively, and were correct in every evaluated field in 54.6% and 19.9% of responses. GPT-5.5 was correct in every field in 48.1%. Across all models, 55.4% of responses were fabricated and 14.9% were correct in every field. Among the five tested models first released in 2026, the corresponding proportions were 35.3% and 31.8%, respectively. Conclusions. Fabrication remained common, and no model was correct in every evaluated bibliographic field in more than 54.6% of responses. Models that identify real papers may still misstate their metadata, so references produced with model assistance require verification before use.