Generative weather models need new evaluation approaches to reveal quality

StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences

Machine Learning

Summary

Synthetic weather data models are often checked by comparing individual weather measurements, but this can hide how well the models capture complex weather patterns. The authors found that calibrating these models inaccurately makes many quality measures look similar, even when the models are very different. They suggest testing on separate time periods without mixing data and looking at whole weather sequences to better judge model performance. Their approach helps spot when models fail to capture seasonal changes or relationships between weather factors.

What this means in practice

  • For climate modelers: Evaluate synthetic weather models with held-out calibration to avoid misleading comparisons and improve detection of realistic weather patterns.
  • For energy grid planners: Use better-evaluated synthetic weather sequences to assess seasonal variability impacts on energy demand forecasting and grid management.

Authors

Mustafa Ozaytac, Ozge Karadag Atas

Abstract

Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.