Discovering hidden symmetries in neural network parameters

Discovering Symmetries in Neural Network Parameter Spaces

Machine Learning

Summary

Neural networks have many internal settings that affect how well they learn and perform. Finding patterns or symmetries in these settings can help understand how the networks work and how to improve them. The authors created a new mathematical method to automatically find these symmetries, even ones not seen before. They tested their method on different types of networks, including popular transformer models used in AI.

What this means in practice

  • For machine learning engineers: Identify hidden structural symmetries in trained models to better analyze and improve training stability and generalization.
  • For neural network architects: Use discovered symmetries in subnetworks to design larger models with predictable and robust parameter transformations.

Authors

Bo Zhao, Nima Dehmamy, Robin Walters, Rose Yu

Abstract

Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.