Training order creates identifiable memory traces in language model weights

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

Machine LearningArtificial IntelligenceComputation and Language

Summary

When training a language model on two different data sources, the order in which the data is presented changes the model's final settings, even with the same data. The authors found that this difference creates a kind of ‘memory’ in the model's parameters that reflects the training order and is linked to specific words. This memory can be identified and even influenced, showing which tokens contribute most to the differences caused by the training order. Their work reveals how subtle interactions during training leave lasting, traceable effects in language models.

What this means in practice

  • For machine learning engineers: Detect and adjust for unintended training order effects on language model outputs by using token-level memory signatures.
  • For ai system auditors: Use the identified memory component to verify the training history or order of model updates in language models.

Authors

John Sweeney

Abstract

Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $η$ on each of sources $A$ and $B$, the weight difference $θ_{AB}-θ_{BA}$ is, to leading order, $η^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $θ_{AB}-θ_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.