Hidden neurons separate stability and storage limit in associative memory networks

Phases in a class of associative memories via hidden neurons

Machine LearningNeural and Evolutionary Computing

Summary

Associative memory networks try to remember patterns by settling into stable states based on inputs. The authors study a specific network design that uses hidden neurons to better understand how it remembers and stores information. They find that one part controls how stable the memories are, while another controls how much can be stored. This separation helps explain why the network behaves differently when handling small or very large amounts of information, and might guide new designs of memory systems.

What this means in practice

  • For machine learning engineers: Design associative memory modules that control memory stability and capacity independently using bipartite layered architectures with hidden neurons.
  • For neuroscience modelers: Model brain-inspired memory systems by separating visible and hidden neuron roles to reproduce different memory retrieval phases under varying load.

Authors

Toshihiro Ota, Masato Taki

Abstract

Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the class $H$, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval. At polynomial load the replica method yields the replica-symmetric phase diagrams and closed-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random-energy-model counting, with paramagnetic, condensed, and frozen phases. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load. The regimes differ in their crosstalk statistics, central-limit at polynomial load and large-deviation at exponential load, and the class $H$ splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians.