Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

2026-08-24Computation and Language

Computation and Language
AI summary

The authors studied two popular ways to measure how easy English texts are to read, called Flesch Reading Ease and Flesch-Kincaid Grade Level. They found that for very long documents, these scores mainly depend on the topics covered rather than how readable the language actually is. Using a statistical model, they showed that topic mixtures can predict these scores quite well, but this doesn't necessarily mean the scores reflect true human reading difficulty. Their results also suggest that factors like genre and style are mixed into the topic signals, so the scores might not capture readability alone.

Flesch Reading EaseFlesch-Kincaid Grade Levelreadability scorestopic modeldocument topic distributiontext readabilitystatistical modelingcorpusBrown corpusBNC (British National Corpus)
Authors
Yo Ehara
Abstract
Flesch Reading Ease (FRE) and the Flesch-Kincaid Grade Level (FKGL) are widely used readability scores for English computed from the same two document statistics, yet their stability on long documents need not imply invariance to lexical composition. Surprisingly, under a topic model with an explicit sentence-boundary token, both scores converge almost surely to deterministic functions of the document topic distribution through just two scalar rates: in the long-text limit, all score variation is mediated by topical composition rather than any residual readability signal. The theory covers both formulae, while the experiments evaluate FKGL. In a fixed admixture with rank[1, q, s] = 3, fibres through interior topic vectors are locally (K-3)-dimensional, whereas regular iso-score level sets are locally (K-2)-dimensional and curved. In out-of-fold evaluation on two balanced corpora, Brown and the written BNC, a topic vector inferred from one document half's content words predicts the other half's FKGL at r = 0.779 and 0.884, respectively. On Brown, adding the topic prediction to genre and mean content-word syllable count yields $ΔR^2$ = 0.002, with a confidence interval spanning zero; on the BNC, the corresponding split-half increment is 0.024, positive in four of five K = 100 fits (median 0.021). Because inferred topics may also absorb genre, register, and style, we do not interpret these results as evidence about human readability or causal effects.