Landau theory of quenched criticality in linear in-context learning

Machine Learning

Summary

The authors study how linear models learn from examples given directly in their input prompt, a process called in-context learning (ICL). They find that when the number of training samples matches the number of model parameters, prediction errors show a sharp spike called double-descent. Using ideas from physics, the authors describe this spike as a phase transition involving fluctuations in the learned parameters. They build a theoretical framework based on Landau theory to explain the behavior and confirm their predictions with numerical experiments. This work links concepts from statistical physics to understanding how linear ICL models behave near critical points.

In-context learningDouble descentLinear modelsStatistical physicsPhase transitionQuenched disorderLandau theoryRidge regressionCritical exponentsSample complexity

Authors

Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han

Abstract

In-context learning (ICL) allows a pretrained model to infer a new task from examples supplied in its prompt without updating its parameters. In linear models of ICL, the prediction error develops a double-descent singularity when the number of pretraining samples becomes comparable to the number of learnable parameters. We formulate this interpolation singularity as a critical phenomenon of a quenched disordered system. By comparing annealed and quenched descriptions of the same linear ICL model, we identify the connected sample-to-sample fluctuations of the learned parameters as the microscopic origin of the singular error. A Landau potential is constructed by integrating the cavity self-consistency equation for the renormalized ridge parameter $ξ$. The role of (magnetization) order parameter is played by $ξ$, while the bare ridge parameter $λ$ becomes its conjugate magnetic field. The normalized sample complexity $τ$ acts as a temperature and the double-descent singularity occurs at the critical temperature $τ_c =1$. The Landau susceptibility is precisely the quantity that diverges in the fluctuation contribution to the prediction error. The order parameter is closely related to the fraction of zero eigenvalues of the empirical relaxation matrix in the ridgeless limit, which define flat directions in the learning dynamics. The Landau theory is generically cubic in the order parameter with critical exponents $(β_{\rm cr},δ_{\rm cr},γ_{\rm cr})=(1,2,1)$. In the large-context regime, there appears a pseudogap-like regime characterized by suppressed order parameter. Predictions of the Landau theory are independently confirmed from numerical solutions of the original learning problem with good quantitative agreement. Our results pave the way for solid statistical-physics understanding of the interpolation criticality in linear in-context learning.