Parameter symmetries link to conservation laws in neural network training
On Parameter Symmetries and Conservation Laws in Gradient Flow
Machine Learning
Summary
Understanding how neural networks learn involves studying patterns in their parameters during training. The paper clarifies when certain symmetrical changes to network parameters lead to conserved values, drawing from ideas similar to those used in physics. The authors create a geometric framework to explain this connection, especially for complex multilayer networks. They also show how this applies to specific network designs like attention models and polynomial neural networks.
What this means in practice
- •For neural network architects: Design multilayer networks with awareness of parameter symmetries that guarantee conserved quantities during training for improved interpretability.
- •For machine learning engineers: Analyze training dynamics in attention-based models to detect invariant properties that might affect network behavior or optimization.
A theory result. No direct application yet.
Authors
Khang Nguyen, Guido Montúfar
Abstract
Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.