Data-driven tuning method improves neural network learning performance
A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks
Machine Learning
Summary
Choosing when to stop training a machine learning model and how big the model should be is a tough problem. The paper presents a new method that automatically decides the best time to stop training and how many 'random features' or neurons to use, based only on the data. This method adapts well to different complexities in the data, works efficiently on large datasets, and applies to neural networks without needing prior knowledge about their settings. It also guarantees strong theoretical learning performance.
What this means in practice
- •For machine learning engineers: Automatically select model size and training time for kernel-based methods on large data to optimize prediction accuracy.
- •For deep learning practitioners: Determine neural network width and halt training dynamically to achieve strong performance without manual tuning.
Authors
Mike Nguyen, Nicole Mücke
Abstract
We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepancy principle, we propose a multi-scale stopping rule that applies to general kernel estimators and show that, unlike previous approaches, it achieves full adaptivity over all smoothness levels in the well-specified case. A key contribution of our work is an extension based on random feature approximations, which reduces computational cost on large datasets while preserving minimax-optimal statistical guarantees. Our procedure not only selects an optimal stopping time but also provides a fully data-driven choice of the number of random features needed to achieve optimal rates. Through the established connection between random features and neural networks in the neural tangent kernel regime, our method further yields a principled, data-driven recommendation for the network width. We prove that the resulting simultaneously chosen width and stopping time allow neural networks to attain minimax-optimal learning rates without prior knowledge of smoothness or capacity parameters.