Deep neural networks reach high accuracy by composing simple functions
Neural Approximation by Function Composition: Rigidity and Doubly Exponential Convergence
Machine LearningInformation Theory
Summary
This paper looks at how deep neural networks use layers that repeatedly apply simple functions to get better at approximating complex shapes. The authors find that certain simple building blocks can only create limited types of curves, like quadratic ones, showing some limits of these methods. Then they design new building blocks that help neural networks approximate functions much faster, with errors shrinking very quickly as networks go deeper. Their work helps explain how the way neural networks are built controls how well and quickly they learn complex functions.
What this means in practice
- •For neural network architects: Guide network depth allocation to achieve faster and more efficient function approximations.
- •For machine learning engineers: Develop more accurate neural models for tasks requiring approximations of smooth functions using improved function composition techniques.
A theory result. No direct application yet.
Authors
Wentao Huang, Haizhang Zhang
Abstract
Deep neural networks approximate functions by composing affine maps with nonlinear activations, but how composition itself creates approximation power is not yet fully understood. We investigate a fundamental mechanism: geometrically weighted sums of iterates of a single scalar generator function. This mechanism underpins the classical tent-map construction of the function \(x - x^2\) and related recursive representations used by Yarotsky, W. E, et al., to analyze the approximation powers of deep neural networks. First, we establish a rigidity theorem: for continuous piecewise linear generators with a finite number of segments, any \(C^3\) function that can be represented in this way is at most quadratic. For non-affine quadratic functions, the geometric factor is at least $1/4$. This result both reveals limitations of the tent-map approach and complements existing methods based on hierarchical bases and recursive polynomial constructions. Second, using an exact remainder identity as guidance, we construct a smooth generator whose iterates yield doubly exponential error decay in total depth for square approximation and, through multiplication modules, for each fixed polynomial. For power series with absolutely summable coefficients on \([-1,1]^d\), distributing depth according to monomial degree yields a uniform approximation error of order \(O(e^{-cL^{1/d}})\) on each interior cube. These findings demonstrate how generator dynamics and remainder estimates govern depth allocation and approximation rates of deep neural networks.