Papers for

neural network architects

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Parameter symmetries link to conservation laws in neural network training

On Parameter Symmetries and Conservation Laws in Gradient Flow

Abstract: Parameter space symmetries and conservation laws play an important role in understanding the loss landscapes and implicit biases of neural networks. Inspired by Noether's theorem in physics, prior works have sought to derive conservation laws under gradient flow from parameter symmetries, but the scope and limitations of this connection remain unclear. We develop a unified geometric framework that clarifies the precise relationship between the two notions, including the conditions under which symmetries correspond to conservation laws. We introduce a notion of compositional identifiability and use it to establish a general inheritance principle for complete characterizations of symmetries and conservation laws in multilayer networks. We apply the framework to multi-head and grouped-query attention, polynomial neural networks, and square deep linear networks.

Mon 28 SeptMachine Learning
The gist
Understanding how neural networks learn involves studying patterns in their parameters during training. The paper clarifies when certain symmetrical changes to network parameters lead to conserved values, drawing from ideas similar to those used in physics. The authors create a geometric framework to explain this connection, especially for complex multilayer networks. They also show how this applies to specific network designs like attention models and polynomial neural networks.
Open → 2609.34549v1

Discovering hidden symmetries in neural network parameters

Discovering Symmetries in Neural Network Parameter Spaces

Abstract: Parameter space symmetries are important for understanding neural networks' loss landscape, training dynamics, and generalization. However, systematically identifying these symmetries remains a challenge. In this paper, we formalize data-dependent parameter symmetries and characterize loss invariance and the group-action axioms through infinitesimal conditions, which provide objectives for jointly learning group generators and nonlinear action maps. Our framework systematically uncovers parameter symmetries, including previously unknown ones. To study larger networks, we establish conditions under which subnetwork symmetries extend to the full model. The same construction gives an explicit family of finite-batch symmetries, providing both analytical examples and a foundation for discovery through small subnetworks. Using the infinitesimal characterization and subnetwork construction, we implement a framework for automated discovery of parameter symmetries, and successfully uncovered symmetries in various architectures, including pretrained transformer models.

Sun 27 SeptMachine Learning
The gist
Neural networks have many internal settings that affect how well they learn and perform. Finding patterns or symmetries in these settings can help understand how the networks work and how to improve them. The authors created a new mathematical method to automatically find these symmetries, even ones not seen before. They tested their method on different types of networks, including popular transformer models used in AI.
Open → 2609.33527v1

Capacity of single neurons and threshold functions precisely quantified

Boolean threshold functions, neuron capacity, and memory retrieval

Abstract: How much information can a single neuron remember? How many memories can neural networks retrieve without creating false memories? These questions are related to a basic question: how many Boolean threshold functions $f(x)=\operatorname{sgn}(a_0+\langle a,x\rangle)$, $x\in\{-1,1\}^n$, are there? In this paper, we show that the number $T_n$ of distinct Boolean threshold functions is \[ T_n=2\binom{2^n-1}{n}\bigl(1+O(n^{-99})\bigr). \] Equivalently, the capacity of a single threshold neuron is $n^2-\log_2(n!)+1+O(n^{-99})$ bits, improving the $O(n)$ error term in the result of Kahn--Komlós--Szemerédi to $O(n^{-99})$. To prove this, we show that, for $1\le r\le n-1$, and $v_1,\ldots,v_r$ are chosen at random from $\{-1,1\}^n$, \[ \mathbb P\!\left\{ \langle v_1,\ldots,v_r\rangle\cap\{-1,1\}^n =\{\pm v_1,\ldots,\pm v_r\} \right\} =1-O(n^{-99}). \] In the context of the Kanter--Sompolinsky Hamiltonian for memory retrieval, this identifies $r=n-1$ as a sharp threshold, at which, for almost every collection of $r$ memories, the only ground states are these memories and their negatives, confirming a weaker form of the Kalai--Linial--Odlyzko conjecture. It also settles a recent open problem posed by M. Anthony on the specification number of Boolean threshold functions. In addition, we show that, for every $1\le r\le n-1$, \[ \mathbb P\{v_1,\ldots,v_r\text{ are linearly dependent}\} =2\binom r2\,2^{-n}+O\!\left(2^{-n}e^{-cn}\right), \] confirming a conjecture of Kahn--Komlós--Szemerédi.

Thu 24 SeptDiscrete MathematicsNeural and Evolutionary Computing
The gist
This paper addresses how much a single neuron can remember and how many memories neural networks can retrieve without errors. The authors precisely count the number of Boolean threshold functions a neuron can implement, improving previous estimates with much smaller error margins. They also find a sharp boundary where collections of memories can be perfectly recalled without false states, confirming long-standing conjectures. Their work improves understanding of neuron memory capacity and helps clarify mathematical properties of neural networks.
Open → 2609.29756v1

Deep neural networks reach high accuracy by composing simple functions

Neural Approximation by Function Composition: Rigidity and Doubly Exponential Convergence

Abstract: Deep neural networks approximate functions by composing affine maps with nonlinear activations, but how composition itself creates approximation power is not yet fully understood. We investigate a fundamental mechanism: geometrically weighted sums of iterates of a single scalar generator function. This mechanism underpins the classical tent-map construction of the function \(x - x^2\) and related recursive representations used by Yarotsky, W. E, et al., to analyze the approximation powers of deep neural networks. First, we establish a rigidity theorem: for continuous piecewise linear generators with a finite number of segments, any \(C^3\) function that can be represented in this way is at most quadratic. For non-affine quadratic functions, the geometric factor is at least $1/4$. This result both reveals limitations of the tent-map approach and complements existing methods based on hierarchical bases and recursive polynomial constructions. Second, using an exact remainder identity as guidance, we construct a smooth generator whose iterates yield doubly exponential error decay in total depth for square approximation and, through multiplication modules, for each fixed polynomial. For power series with absolutely summable coefficients on \([-1,1]^d\), distributing depth according to monomial degree yields a uniform approximation error of order \(O(e^{-cL^{1/d}})\) on each interior cube. These findings demonstrate how generator dynamics and remainder estimates govern depth allocation and approximation rates of deep neural networks.

Tue 22 SeptMachine LearningInformation Theory
The gist
This paper looks at how deep neural networks use layers that repeatedly apply simple functions to get better at approximating complex shapes. The authors find that certain simple building blocks can only create limited types of curves, like quadratic ones, showing some limits of these methods. Then they design new building blocks that help neural networks approximate functions much faster, with errors shrinking very quickly as networks go deeper. Their work helps explain how the way neural networks are built controls how well and quickly they learn complex functions.
Open → 2609.25874v1

Neural dynamics must respect language structure for real comprehension

Formal Properties of Language as Constraints on Neural Dynamics

Abstract: What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural activity. Yet predictive success leaves mechanisms under-constrained. Here, we show that algebraic properties of language specify invariants that mechanisms must preserve: non-associative hierarchical grouping, commutativity, recursive closure, access to substructures, and structured workspace transitions. We term this the Neural Admissibility Program (NAP). Syntactic structure building is analyzed algebraically, with candidate mechanisms offered for each requirement: content-addressable workspace memory, graph-structured transient dynamics scheduling structure-building operations (e.g. stable heteroclinic channels), and a phase-coupled sealing operation recording grouping. Simulations show that a corrected Marcolli-Berwick entropy-optimized binding gate preserves grouping only within a narrow commitment band. As an alternative, we propose a novel neural binding operation we term 'Meld': two constituent populations converge through shared synapses, integrate sublinearly, and saturate. Meld is, to our knowledge, the closest neurally plausible composition law to syntactic Merge. It preserves every NAP invariant, uses known cortical operations, and recovers hierarchical structure at every tested depth and temperature. It predicts that effective population dimensionality separates alternative bracketings and that the composite depends on constituent disagreement. Importantly, the laws decoding bracketing most accurately are a priori inadmissible, showing that decoding accuracy alone cannot adjudicate between mechanisms. By specifying how neural dynamics can remain faithful to linguistic structure, the NAP changes the criterion by which neural implementations of cognition are evaluated.

Sun 13 SeptComputation and Language
The gist
Understanding language requires certain basic rules to be followed by brain processes, such as keeping words grouped properly and recognizing repeated patterns. The authors propose a formal framework called Neural Admissibility Program (NAP) that spells out these essential rules as mathematical properties. They test different neural mechanisms and introduce a new one called Meld, which closely matches how language structures work in the brain. Their work shows that just measuring how well a model predicts brain data is not enough to understand if it truly captures language processing.
Open → 2609.14384v1