AI summaryⓘ
The authors study how the common optimization method called gradient descent (GD) behaves and differs when viewed continuously versus as discrete steps, especially when using hard-ReLU activation functions. They show that while the discrete gradient descent steps and their exact derivatives converge to a certain continuous description, this continuous limit misses some key event-related behaviors in the derivative. Their work carefully separates smooth changes from sharp jumps caused by activation events and explores conditions under which these differences matter, particularly in convex settings. The results focus on controlled, finite-time scenarios and do not claim to describe behavior in large-scale training.
gradient descentgradient flowhard-ReLUautomatic differentiationactivation eventsStieltjes representationconvexityHessianadjointsfinite-horizon dynamics
Abstract
Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. We prove that, over a fixed finite horizon, the GD states converge and these exact discrete derivatives approach an event-free regional propagator, whereas the derivative of the limiting flow also contains speed-normalized activation-event transfers. A prepoint Stieltjes representation separates the absolutely continuous regional Hessian from atomic interface curvature; one nonzero gradient jump produces an exactly rank-one endpoint discrepancy, and global convexity prevents complete multi-event cancellation whenever an event is strict. Nevertheless, a standard family of globally 1-strongly convex residual-ReLU squared-loss risks realizes arbitrarily large reciprocal sensitivity ratios on open initialization sets, with a uniform transversality margin. The same discrete-versus-flow decomposition extends to parameters and reverse-mode adjoints; resolved smoothing in the scalar or autonomous-normal regime and consistent event localization recover the flow sensitivity. The results concern deterministic full-batch, finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events; they are consistency theorems, not prevalence claims for large-scale training.