Papers for

climate model developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Extreme event prediction improves with local instability sensing

Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

Abstract: Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only \(50\) time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with $20$ times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.

Fri 25 SeptMachine Learning
The gist
Predicting rare, extreme events in chaotic systems is tough because these events depend on short-term instabilities that are not common in usual data. The authors developed a method that uses a group of slightly different simulations to detect these local instabilities without complex calculations. By feeding this information into machine learning models, the method improves the prediction of rare events, even with limited training data. They tested their approach on both simple chaotic systems and more complex fluid models, showing better accuracy compared to traditional methods.
Open → 2609.30746v1

Contrastive learning clarifies cloud model differences and observations

Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach

Abstract: Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observations. We develop an explainable contrastive learning model that maps 5 monthly cloud and radiation fields into a shared representation space. We train the model on the fields of two 100-member Community Atmosphere Model version 6 (CAM6) PPEs, spanning 34 parameters, that only differ in the warm rain microphysics scheme: KK2000, the default bulk microphysics scheme, and TAU-ML, a neural network emulator of a bin microphysics scheme. The learned representations separates two PPEs with over 94\% linear classification accuracy while preserving the seasonal variability and ensemble spread due to parameter perturbations. In the shared representation space, the representations of satellite observations occupy the same low-dimensional manifold as the PPEs but are displaced from them most strongly during boreal spring and autumn. TAU-ML PPE has a lower distance to observations compared to KK2000 in the representation space. Integrated Gradients attributions highlights the contributions in subtropical low-cloud regions, Northern and Southern Hemisphere storm track regions, and tropical convection regions to differences between PPEs and observations. Regional attributions correlate most strongly with parameters associated with cloud microphysics, boundary layer turbulence, and deep convection. These results demonstrate that explainable representations of climate fields can attribute model differences to specific variables, regions, seasons, and physical parameters.

Thu 24 SeptArtificial Intelligence
The gist
Climate models have many settings that affect their weather predictions, but it is hard to understand how these settings influence complex weather features. The authors used a machine learning tool to create simple summaries of many climate patterns, which helps to compare two versions of a climate model and real satellite data. Their approach shows which regions, seasons, and weather variables differ between models and observations, linking these to specific model settings. This helps scientists see where models match reality and where they need improvement.
Open → 2609.30420v1

Hybrid AI models improve weather forecasts using learned memory variables

Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation

Abstract: Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which the subgrid state depends only on the current coarse state with no memory of previous states, which is unrealistic for processes such as convection that have intrinsic persistence. To address this, we enhance local-in-time parameterizations by learning prognostic variables that compactly carry important, additional past information where no explicit sub-grid information is available. First we compress past information into a low-dimensional latent space using an autoencoder, which then informs a neural network trained to parameterize targeted subgrid-scale processes. We then replace the autoencoder with symbolic equations that govern the time evolution of the latent variables, yielding additional prognostic memory variables that can be integrated alongside the resolved atmospheric state. We evaluate this approach on two systems: the Lorenz-96 model (online) and surface precipitation from high-resolution atmospheric simulations (offline). A forced multivariate linear ordinary differential equation recovers most of the added value achieved by the autoencoder-based approach in both experiments. Benchmarked against diagnostic parameterizations without memory, our memory-informed approach improves climate statistics and temporal structure, including a realistic diurnal cycle of tropical land precipitation.

Mon 21 SeptMachine Learning
The gist
Weather and climate models often miss important slower processes because they only look at the current state and ignore past information. This paper shows how to add a kind of memory to AI models that predict weather patterns by learning compact variables that capture past information. They use math equations to describe how these memory variables change over time, making forecasts more realistic. Their method improves predictions of rain patterns and the daily cycle of tropical weather compared to models without memory.
Open → 2609.24882v1

Machine learning weather models struggle with energy flow and error growth patterns

Butterfly Effect and the Kinetic Energy Cascade in Probabilistic Machine Learning Weather Prediction Models

Abstract: This study analyses kinetic energy (KE) spectra, difference kinetic energy (DKE) spectra, and signatures of KE transfer across spatial scales in four state-of-the-art probabilistic machine learning weather prediction (MLWP) models: NeuralGCM-ENS, FourCastNet 3, AIFS-ENS, and GenCast. Results are compared with those from the physics-based numerical weather prediction model IFS-ENS. While NeuralGCM-ENS successfully reproduces the expected upscale transfer of KE, noise injection at its encoder stage underestimates mesoscale KE. Conversely, AIFS-ENS, GenCast, and FourCastNet 3 produce realistic KE spectral magnitudes but do not capture the expected upscale transfer of KE. In particular, AIFS-ENS and GenCast, which employ spatially uncorrelated stochastic perturbations, exhibit enhanced accumulation of KE at high wavenumbers. All examined models exhibit upscale error growth, reflected by the progressive shift of the DKE spectral peak toward larger wavelengths over time. However, the MLWP models struggle to reproduce the rapid initial growth of ensemble spread at small spatial scales associated with the butterfly effect. The results show that MLWP models can misrepresent the known scale transfer of kinetic energy despite producing skilful weather forecasts.

Wed 16 SeptMachine Learning
The gist
The study looks at how well several advanced machine learning weather models handle the flow of energy at different scales in the atmosphere and how errors spread out over time. The authors find that while some models mimic certain energy patterns well, others do not capture important ways energy moves between scales. All models have trouble showing the quick early spread of small errors known as the butterfly effect. This means that even though the machine learning models can predict weather well, they don't fully represent how energy and errors behave in the real atmosphere.
Open → 2609.18489v1

Physics-informed model predicts Mars nightside atmospheric gases reliably

Physics-Informed Multi-Task Surrogate Model for the Martian Nightside Thermosphere

Abstract: Modeling the Martian nightside thermosphere remains challenging due to sparse in situ sampling and strong coupling among transport, magnetic, and seasonal processes. Purely data-driven models can produce non-physical artifacts, such as density inversions, in poorly sampled altitude regimes. We present a multi-task physics-informed neural network that simultaneously predicts the base-10 logarithmic densities of four neutral species (O, CO$_2$, N$_2$, and Ar) using more than a decade of MAVEN/NGIMS observations (MY 32-38, 2014-2025). A shared backbone learns a common representation of the nightside thermospheric state and branches into species-specific output heads. A weak monotonicity prior is incorporated via automatic differentiation by penalizing positive vertical gradients in logarithmic density. Experiments using an orbit-disjoint train/validation/test split show that physics-informed regularization substantially reduces non-physical inversions while preserving predictive skill and slightly improving it in the best-performing configuration, as measured by RMSE, MAE, and $R^2$. The resulting model provides a computationally efficient surrogate for nightside thermospheric reconstruction with improved vertical consistency.

Wed 9 SeptMachine Learning
The gist
Understanding the atmosphere on Mars's night side is hard because there aren’t many direct measurements and many processes are linked. Purely data-based methods sometimes make unrealistic predictions like densities that don’t decrease with altitude. The authors created a neural network that predicts the amounts of four gases at once, using over ten years of spacecraft data. They included physics rules to stop impossible results and found it improved the model’s realism without losing accuracy. This model can quickly recreate the state of Mars’s nightside upper atmosphere with better vertical consistency.
Open → 2609.10077v1