New framework reveals hidden flaws in weather model dynamics

Evaluating Dynamical Fidelity through Predictive Structure in Physical Representations

Machine Learning

Summary

Weather prediction models are usually judged by how close their forecasts match real weather and if they obey known physical rules. But these checks don't always show whether the models truly capture how weather changes over time. The authors created a method that uses expert-chosen ways of representing weather processes to test if models follow real dynamic patterns. They tested this on three advanced weather models and found differences that normal error tests missed. This helps experts better understand where models fail and how to improve them.

What this means in practice

  • For weather forecasters: Identify weaknesses in weather models by testing if they reproduce real atmospheric dynamic patterns beyond just forecast errors.
  • For climate model developers: Use expert-defined representations to systematically evaluate and improve the physical realism of climate simulations.

Authors

Oskar Bohn Lassen, Joao Paulo de Souza Boger, Simon Driscoll, Stephen I. Thomson, Sebastian Schemm, Filipe Rodrigues, Francisco C. Pereira

Abstract

Machine-learning models for physical systems are currently evaluated primarily through errors between predicted and reference states and, increasingly, through tests of physical consistency. These metrics assess whether predictions are accurate and satisfy selected physical requirements, but provide limited insight into whether learned trajectories reproduce the underlying dynamics. Domain experts examine such relationships through physical representations that expose relevant processes, interactions, and responses, but these analyses are often separated from typical machine-learning evaluation. We introduce a practical framework for evaluating dynamical fidelity through predictive structure in physical representation spaces. Experts define the representations, while reference trajectories determine which relationships are predictive and retained as evaluation tests. We demonstrate the approach in atmospheric forecasting using ERA5 representations of planetary-wave activity and Northern Annular Mode evolution, and evaluate Pangu-Weather, GraphCast, and FengWu. The models exhibit distinct departures from reference predictive structure that are not reflected by conventional forecast errors. The framework thereby turns domain-expert representations into systematic tests of learned physical dynamics without prescribing the relationships in advance.