Models organize reusable functions by linking parameters to components

The Ups and Downs of Backprop Weights

Machine LearningArtificial Intelligence

Summary

Deep learning models learn by adjusting many parameters, but sometimes these parameters overlap across tasks, making it hard to change one function without affecting another. The authors identify this problem as weight entanglement and propose a way to organize parameters into reusable modules called weight operators. Their approach lets models select and update only the relevant modules for each task, keeping functions more separate and adaptable. This could help make learned skills easier to reuse and modify independently in future AI systems.

What this means in practice

  • For machine learning engineers: Design deep learning models that selectively update reusable parameter modules, enabling more efficient multi-task learning and adaptation.
  • For software developers in ai: Build AI systems that reuse learned functions systematically by linking parameters to specific functional components, improving modularity.

Authors

Giuseppe Chindemi, Benjamin F. Grewe

Abstract

Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large hierarchical networks to learn complex functions end-to-end. Yet it does not by itself determine how parameters should be organized so that functional components can be reused and adapted selectively. For example, object recognition and motion prediction may depend on overlapping parameter sets, making them difficult to isolate or modify independently. We call this condition weight entanglement. Modern architectures dynamically select which parts of a network process each sample: nonlinearities gate units, attention selects interactions, and Mixture-of-Experts architectures route inputs to modules. Yet such selection does not ensure that the same functional component remains linked to an identifiable parameter set across samples. We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample. Learning proceeds in two stages: the model first infers the required operator composition, then updates only the selected operators' parameter sets. Vector Networks (VNs) provide one implementation. They couple operator selection to local error-driven updates within each layer and show that learned operators can be reused in combinations absent from training while updates remain restricted to the selected parameter sets. This provides a basis for testing functional parameter identifiability: whether an operator remains linked to the same functional component during learning. We argue that functional parameter identifiability may provide an organizing principle for models that systematically reuse and recombine learned functions while adapting only the components that need to change.