Large language models show limits in mimicking infant syntax learning

A retrospective analysis on the use of LLMs to study infant syntax learning

Computation and Language

Summary

This paper looks back at how large language models (LLMs) have been used to study how babies learn sentence structure. The authors point out that the way these models are trained and tested involves a lot of assumptions that limit what we can conclude. They also find that using baby-like learning data doesn’t notably improve model performance on usual language tests. This suggests that actual infant learning processes might be quite different from how LLMs work.

What this means in practice

  • For language technology developers: Avoid relying solely on infant-like training data for improving syntax capabilities in language models based on BabyLM challenge insights.
  • For computational linguists: Refine experimental designs by critically assessing assumptions in infant syntax learning models informed by BabyLM challenge methodology.

A position paper. It proposes an approach and reports no results.

Authors

Hélie Bazin, Anouk Barberousse, François Yvon

Abstract

Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.