Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models

2026-07-12Machine Learning

Machine LearningArtificial Intelligence
AI summary

The authors focus on improving how deep generative models simulate rare, risky events that are important for decision-making when data is limited. They point out that usual training methods prioritize common cases, making the rare events less reliable. To fix this, they introduce Diachronic Sample Integration (DSI), which combines samples from multiple training checkpoints instead of depending on just one. Their tests show that DSI better estimates rare events without changing how the model is trained, making it more stable for risk-sensitive uses.

Deep generative modelsRare event simulationTail distributionCheckpoint ensembleBias-variance tradeoffStochastic trainingSimulation budgetRisk-sensitive applicationsMultivariate processesHigh-frequency trading
Authors
Shuning Zhao, Patrick Wong, Leran Zhang, Xiaolin Hu
Abstract
Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applications their usefulness depends on rare adverse scenarios rather than typical samples. Standard generative objectives prioritize bulk distributional fidelity, leaving low-probability tails vulnerable to localized optimization noise and making tail-dependent functionals unstable under finite simulation budgets. We introduce Diachronic Sample Integration (DSI), a test-time inference framework that ensembles generated samples across checkpoints from a stochastic training trajectory. DSI targets a checkpoint-mixture distribution that averages checkpoint-specific tail fluctuations rather than relying on a single brittle endpoint. We formalize this mechanism through a finite-budget bias-variance theory. Empirically, across multivariate synthetic processes and high-frequency trading data, DSI substantially reduces tail-estimation error compared to single-checkpoint baselines under fixed simulation budgets, outperforming standard diffusion and state-of-the-art tail-aware baselines without modifying the generative objective.