Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample

2026-08-31Machine Learning

Machine Learning
AI summary

The authors developed a method called Flow approximate leave-one-out (Flow-ALO) to estimate how well a specific type of learning process (a smooth nonconvex gradient flow) performs on new data by approximating the effect of leaving out one training example. They carefully break down the errors in their estimation and prove precise bounds on these errors under certain mathematical conditions. Their approach works even when some usual matrix properties don't hold and applies to neural networks with two layers, maintaining accuracy regardless of the network size. Overall, the authors provide a controlled way to understand the population risk curve without needing direct access to all data variations.

conditional population-risk curvenonconvex gradient flowFlow-ALOleave-one-out (LOO) approximationHessian matrixmean-field networksdeletion-response errorjackknifelocal Lipschitz continuitysmooth two-layer neural networks
Authors
Mingzhi Song
Abstract
We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO fluctuation, and deletion-to-full risk transfer. On each fixed finite horizon, bounded centered training-loss gradients, a one-sided Hessian lower bound, locally Lipschitz Hessians, and a strict tube-closure condition yield an explicit $(n-1)^{-2}$ bound for the deletion-response error. Bounded evaluation-loss gradients transfer the deletion-response bound to the score without requiring the Hessian to be invertible. Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve. For bounded smooth two-layer mean-field networks training both layers, the score-error bound is uniform in width.