Memorization happens mostly in lower layers of language models during training

Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models

Machine LearningComputation and Language

Summary

Language models sometimes memorize parts of their training data, but scientists haven’t fully understood how and when this happens during training. The authors looked closely at how models learn both repeated sequences and rare sequences from their training data by studying gradients during training. They discovered that memorization often occurs in the lower layers of the model and that memorized sequences can be forgotten or reinforced depending on how they align with other training signals. They also showed it’s possible to predict and even reduce memorization by focusing on key model parameters.

What this means in practice

  • For language model developers: Identify and reduce memorization in large language models by monitoring gradients and intervening on key lower-layer parameters.
  • For machine learning engineers: Predict which training sequences are memorized early in training to guide dataset curation and augmentation strategies.

Authors

Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank, Michael A. Hedderich

Abstract

Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.