Data-efficient language modeling improves prediction with less text
Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
Computation and LanguageArtificial Intelligence
Summary
Training language models with limited text is challenging because they must understand context, handle new inputs, and remember useful information. The authors conducted a multi-stage research program using a small amount of text and word presentations to build better models. They discovered that organizing training around specific context relationships and carefully managing what the model sees and learns helps improve performance. Their improved models achieved the best results on a public small-data benchmark and are available for others to use and study.
What this means in practice
- •For natural language processing engineers: Develop language models that deliver higher accuracy using much less training text by following proven data-efficient training principles.
- •For language technology product teams: Improve customer-facing language understanding features in low-data scenarios by applying models validated on strict small-data benchmarks.
Authors
Shuxing Yang, Kaihao Zhu, Junjie Yang, Rui Zhao, Junyao Wu, Yize Wang, Wenhao Li, Fujia Chen, Taowen Deng, Shenzhan Hong, Yaqi Li, Zichen Li, Jincheng Mi, Yuang Pan, Hongsheng Chen, Yihao Yang
Abstract
Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations. Three stages connected frontier advancement, principle discovery, and principle-guided model improvement. Stage I combined compact restatements, budget reinvestment, and residual incremental learning to build a frontier model. Stage II found that exact repetition and aligned restatement produce different patterns of context use, depending on target relations and prediction windows. In controlled tasks, recovering familiar performance did not ensure that unseen inputs could still use learned computations. These findings support a testable data-efficient learning principle: organize experience around the contextual dependencies needed for prediction; separately design visible information, supervision, and preservation; test learning, generalization, and retention. Stage III retained source text, masked more local clues, supervised selected targets, and preserved predictions on ordinarily masked inputs. Two continuation seeds from the same parent outperformed ordinary continuation on the complete nine-metric aggregate. Overall rose from 42.02 to 42.25 across two generations; the second achieved the highest Overall in the public Strict-Small snapshot of 8 September 2026. Further studies addressed compression, relational anchors, shared representations, and measurement. Models are available on Hugging Face; code and research records accompany the GitHub repository. Together, these stages illustrate Research RSI: recursive self-improvement of the research process. Scientific understanding and method innovations change subsequent questions and designs; new experiments test and refine them.