More than half of recent astronomy papers use language models
More than half of recent astronomy papers are written with language-model assistance
Computation and LanguageDigital Libraries
Summary
It is hard to know how often scientists use language models to help write papers. The authors looked at over 200,000 astronomy papers from 2015 to 2026 and found a hidden pattern in the words that shows when language models were likely used. They estimate that over half of recent papers have some language-model assistance, even though very few authors openly say so. They also show that as authors learn to avoid words that reveal help, it becomes more difficult to spot language-model use over time.
What this means in practice
- •For academic journal editors: Estimate how often language models assist writing without relying on self-reports, to adapt editorial policies or peer review guidelines.
- •For scientific publishers: Develop automated tools to detect subtle language-model assistance in submitted manuscripts by using distinctive vocabulary patterns.
Authors
Serat M. Saad, Yuan-Sen Ting
Abstract
Language models leave a distinctive vocabulary in the prose they help write, and we measure how much of the astronomy literature now carries it. From the full text of 207,111 astro-ph papers spanning 2015 to mid-2026, we count those words in each paper and model the counts, in proportion to paper length, as a mixture of assisted and unassisted writing in a hierarchical Bayesian model. Papers from before 2020 calibrate the unassisted rate, and the 392 papers that disclose model use calibrate the assisted one. Our answer depends on how often these words would appear today if nobody used a model, a rate that must be modeled rather than observed, so we extend it past 2020 under three assumptions and report all three. For 2025 that gives $54^{+8}_{-8}\,(\mathrm{stat},\,95\%)\,^{+26}_{-0}\,(\mathrm{sys,\ background})$% of papers, the second error being the spread across the three. The estimate stays at or above 36% when we vary that choice, the calibration, and the requirement that adoption only rises. A word list built from the astro-ph corpus, keeping only words that rose across every subfield, leaves 2025 in the same range. Assisted writing is also getting harder to see, since authors adapt to the words that reveal it and the marker excess more than halves between 2023 and 2026. Our model allows for that fading, so it can separate a fainter trace from reduced use. More than half of recent astro-ph papers therefore carry a language-model trace, while only 0.81% of 2025 papers disclose it, one declaration for every $\sim$66 papers with a trace.