Spectral regularization improves self supervised learning representations

$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

Machine LearningComputer Vision and Pattern Recognition

Summary

Self-supervised learning tries to teach computers to understand images or videos without labels by creating multiple views of the same data and making their features similar. However, sometimes the learned features can become too simple or collapsed, limiting how well they work on other tasks. The authors introduce a new math-based technique called SACReg that helps keep the features diverse and rich by balancing layer weights, which prevents collapse. They used this idea to improve existing methods, showing better performance on image and video recognition tasks.

What this means in practice

  • For computer vision engineers: Enhance image classification models by increasing representation quality through spectral anti-collapse regularization in self-supervised training.
  • For video analytics developers: Improve video understanding models on benchmarks like Something-Something-v2 using λ-JEPA to obtain better learned features without labeled data.

Authors

Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello

Abstract

Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $λ$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $λ$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $λ$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $λ$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $λ$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.