Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries
2026-07-27 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors study a phenomenon called grokking, where models suddenly generalize well after a delay during training. They analyze linear models with weight decay and show that slow improvement happens in a special subspace where training outputs don't change, controlled only by weight decay. Their theory matches known scaling laws for grokking time and predicts how different optimizer settings affect this process. They confirm their predictions exactly in a simple model and also see similar behavior in a modular addition task.
grokkingweight decayheavy-ball optimizationlinear modelsnull spacepopulation riskregularizationlate-time relaxationmodular additionoptimizer
Authors
Taeyoung Kim
Abstract
Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-β)/(ηλ)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.