Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework
2026-08-31 • Machine Learning
Machine LearningArtificial Intelligence
AI summaryⓘ
The authors developed Topo², a new way to separate and measure how deep neural networks both learn true patterns and memorize wrong labels when trained on noisy data. They use a topological tool called persistent homology to split the network's internal representations into two parts: one that captures real class structure and another that reflects memorized errors. Their method lets them control memorization and generalization independently and defines precise laws governing this behavior. They also show what doesn’t explain the patterns and turn memorization from a vague concept into something quantifiable and removable.
deep neural networksnoisy labelsmemorizationgeneralizationpersistent homologytopological data analysisrepresentation spacemanifoldCIFAR-10SVHN
Authors
Zhanbo Zhang, Ming Liu, Qing Wang
Abstract
Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed. Persistent-homology H1 structure of the representation space separates into a within-class manifold channel (a function of the training stopping point) and a cross-class channel (a monotone readout of memorized flipped samples). An intervention, the FM0 prescription (zero loss on flipped samples from epoch 0), reaches each setting's generalization ceiling while memorizing essentially nothing. Within the framework we establish a law set with graded evidence: (L2) FM0 separation prescription (9/9); (L1) the within-channel as a training-position function (mid-rise 6/6; convergence-back CIFAR 3/3, SVHN 2/3); (L3) a ring-construction identity (definitional, not a law); and TLS (memory-generalization topological layering): memory is causally additive, anchored (silencing clean collapses the representation), invertible (stripping memory restores near-ceiling generalization), and quantitatively billable (the memorization cost law, effective slope coefficient C ~ 0.38 at the reference capacity: CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715, capacity-dependent in general and traced to clean-sample feature displacement). We also publish the framework's boundaries: a falsification ledger of nine dead ends, and an instrument-vindication section that excludes six families of global statistics as explanations of the within-channel. The framework turns "memorization" from an ill-defined capacity into a measurable, separable, invertible topological layer.