Learning with weighted data improves models on long memory sequences

Weighted Empirical Risk Minimization for Machine Learning under Long-Range Dependence: Exact Pathwise Rates and Learning-Error Geometry

Machine Learning

Summary

This paper studies how machine learning models perform when trained on data that has long-range dependencies, which means past data points affect future ones over a long time. The authors show precise mathematical results about how weighted training data influences the learning process and the accuracy of the models. They describe how certain technical properties of the data and weighting affect learning speed and error behavior. Their findings help understand the best way to weight data in time series prediction and classification when data is strongly correlated over time.

What this means in practice

  • For time series analysts: Improve model accuracy for forecasting by choosing better weighting schemes in correlated time series data.
  • For data science teams: Understand the limits and error behavior of weighted learning methods when data shows long memory effects.

Tested on simulated data.

Authors

Elina Moldavskaya

Abstract

We develop an exact almost-sure learning theory for smooth parametric models trained by regularly weighted empirical risk minimization on long-range dependent data. The training observations are generated from a fixed finite window of a stationary Gaussian sequence, and the sample weights are regularly varying. If the loss gradient at the population minimizer has Wiener-chaos rank $m$ and a nonzero low-frequency coefficient, then, in the long-memory interior regime, the finite-lag score reduces on the iterated-logarithm scale to a single weighted Hermite chaos. This yields an almost-sure Bahadur representation, an exact limsup law for the learned parameter, and, for $m\ge2$, the functional cluster set of the complete learning trajectory. The polynomial learning exponent is determined by the memory parameter and the chaos rank and is invariant under the admissible power weighting, whereas the sharp pathwise constant and cluster geometry depend on the weights. In the rank-one case, global optimization over the admissible power exponents shows that every optimizer is positive. Time-series prediction and classification examples illustrate the results.