Korean academic language shows rising AI influence from 2024 to 2026

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Computation and LanguageDigital Libraries

Summary

This study looks at how the words used in Korean academic article summaries changed after 2022. The researchers found that from late 2024 to 2026, certain words related to AI-generated or influenced text became much more common. For example, the word for 'suggest' appeared four times more often than expected. The change is not due to translation or subject differences, and similar trends appear in English summaries about a year earlier. This suggests that large language models (LLMs), like AI tools, are increasingly shaping academic writing in Korean.

excess vocabularyKorean Journal abstractsmorphological unitslarge language models (LLMs)academic writingword frequency trendstranslation effectssubject-matter controllanguage model annotationKorean Corpus Index (KCI)

Authors

Aron Lee

Abstract

Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018-August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1-2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: sisahada "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like araboda "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024-2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025-2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent.