The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning

2026-08-03Artificial Intelligence

Artificial IntelligenceMachine Learning
AI summary

The authors examine why machine learning models sometimes seem limited in predicting labels that summarize long time periods but only use short time-window inputs. They show that this limit often comes from how and when data is collected, rather than the model's ability. By breaking down variability in the labels into stable traits and changing states, they explain why one short snapshot predicts general differences well but misses within-person changes over time. Their work helps separate what limits come from model design versus data collection methods and highlights that the type of label determines the important timescale, not just the amount or timing of data.

Bayes risklatent Gaussian processtest-retest reliabilitystate-trait decompositiontemporal correlationmachine learning benchmarklabel varianceacquisition protocoloccupation timeeffective temporal span
Authors
Xizhe Zhang
Abstract
Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $Θ_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.