Papers for

climate modelers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Cascade model architecture reduces errors in long term neural predictions

CasEm: A Cascade Architecture for Long-Horizon Neural Emulation

Abstract: Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolving model of physically specified aggregates. Its forecasts guide corrections to full-state predictions, without feedback from the backbone to the aggregate model. Effective guidance requires aggregates that cover substantial backbone error, remain accurately predictable, and support useful full-state corrections. We derive a finite-horizon error bound that clarifies these three factors and use empirical diagnostics to guide subsystem selection. Across four ODE/PDE benchmarks, CasEm reduces long-horizon rollout errors across diverse backbones and suppresses the trend toward error divergence in both diffusion tasks using Fourier neural operator backbones. In global climate emulation, CasEm with a regional total-water subsystem reduces 10-year full-state time-mean error by 66.6% and 46.3% for frozen ACE and Spherical DYffusion backbones, respectively, while adding less than 3% to inference time.

Mon 28 SeptMachine Learning
The gist
Neural networks that predict future states over long periods can make bigger and bigger mistakes as they go. The authors introduce a new method called CasEm, which adds a second model focusing on key summary features to help correct these errors during long predictions. This second model runs separately and guides corrections without interfering with the main model’s short-term accuracy. Their approach cuts down mistakes in various physics simulations and climate models, making long-term forecasts more reliable.
Open → 2609.34246v1

Generative weather models need new evaluation approaches to reveal quality

StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences

Abstract: Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.

Sun 27 SeptMachine Learning
The gist
Synthetic weather data models are often checked by comparing individual weather measurements, but this can hide how well the models capture complex weather patterns. The authors found that calibrating these models inaccurately makes many quality measures look similar, even when the models are very different. They suggest testing on separate time periods without mixing data and looking at whole weather sequences to better judge model performance. Their approach helps spot when models fail to capture seasonal changes or relationships between weather factors.
Open → 2609.33761v1

Aurora-X improves forecasting accuracy across many time series types

Aurora-X: Built for Extreme Time Series Forecasting

Abstract: Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon lengths, and future covariates if available during midtraining. Variable-resolution post-training further enables an adjustable temporal span per token at inference. With fixed model weights, this supports longer histories under a fixed token budget or fewer tokens for the same history, enabling test-time scaling. With a versatile architecture, Aurora-X supports cross-variable modeling, covariate conditioning, and parallel decoding of future patches for probabilistic forecasting. These are supported by a novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series. Furthermore, we propose an implicit quantile network head that predicts arbitrary quantiles to characterize predictive distributions, enhancing probabilistic forecasting flexibility. Comprehensive experiments on GIFT-Eval, TIME, FEV-Bench, TFB, and DAG-Bench demonstrate state-of-the-art forecasting performance against pretrained TSFMs and task-specific supervised models.

Fri 25 SeptMachine Learning
The gist
Time series forecasting tries to predict what will happen next based on past data, like weather or stock prices. The authors introduced Aurora-X, a very large model designed to handle many different types of time series data better than before. It learns patterns step-by-step, including relationships between variables and different time lengths, and it can adapt to using longer or shorter past data when making predictions. Aurora-X also predicts probabilities of future values, providing a fuller picture of uncertainty. Tests show it usually forecasts more accurately than earlier models.
Open → 2609.31038v1

Machine learning models predict climate effects with causal insights

Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models

Abstract: Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate model. As a key advance over previous work, our framework explicitly models both atmospheric dynamical interactions arising from internal climate variability and forced responses due to changes in atmospheric greenhouse gas and aerosol concentrations. When trained on future climate change scenarios, our method accurately predicts the long-term global mean and regional temperature evolution and shows physically realistic responses to perturbations in greenhouse gas and aerosol concentrations when evaluated on unseen scenarios. Our results underline the potential of causal representation learning frameworks for advancing climate model emulation.

Fri 25 SeptMachine Learning
The gist
Climate models are complex and slow to run, so people use machine learning to simulate future temperature changes faster. The authors developed a new machine learning method that not only predicts temperature changes but also explains how greenhouse gases and aerosols cause these changes. Their method works well on data it hasn’t seen before and matches expected physical behaviors. This approach could help improve trust and understanding of climate model predictions.
Open → 2609.30995v1

Lightweight models improve finer scale climate data prediction accuracy

Lightweight Probabilistic Downscaling from a Deterministic Base Model

Abstract: Climate data downscaling is the task of increasing the spatial resolution of climate data, typically by generating fine-resolution regional climate data from coarse global model output. Recent machine learning (ML) work in the related task of weather forecasting has seen significant improvements due to newly devised training methods and architectural components, but these have not yet benefited downscaling. We adapt two of these methods to create a family of lightweight probabilistic ML downscaling models built on a modified U-Net backbone and evaluate them on the CORDEX-ML-Bench suite for daily maximum temperature and precipitation across three geographic regions: the Alps, New Zealand and South Africa. We find that a two-stage training curriculum, combining deterministic pretraining with probabilistic tuning, transfers well to downscaling, beating the state-of-the-art for RMSE. Our work provides an advancement towards lightweight, probabilistic downscaling models, reducing the current trade-off between computational intensity and distributional fit.

Thu 24 SeptMachine Learning
The gist
Climate models often predict weather and climate at a large scale, but it's useful to have detailed local predictions too. The authors developed smaller, efficient machine learning models to increase the detail of climate predictions, working from coarse global data. They blended two training steps—starting with simple prediction and then tuning for uncertainty—to get better, more accurate results across regions like the Alps and South Africa. Their approach balances accuracy and computational cost better than previous methods.
Open → 2609.29383v1

Functional dynamic mode decomposition learns infinite dimensional systems from data

Functional dynamic mode decomposition: Learning infinite-dimensional systems from data

Abstract: Dynamic mode decomposition (DMD) is a data-driven method that computes the best linear approximation of the underlying dynamical system and decomposes the dynamics into a superposition of characteristic spatiotemporal patterns. Originally introduced by the fluid dynamics community, DMD and its extensions have found widespread use in many other research areas such as molecular dynamics, climate science, engineering, finance, and neuroscience. Applications include dimensionality reduction, forecasting, system identification, control, and spectral clustering. In order to apply DMD to partial differential equations, the spatial domain is typically first discretized using finite difference or finite element techniques, thus implicitly rendering the problem finite-dimensional. We extend projected and exact DMD to infinite-dimensional systems. Rather than estimating matrices from vector-valued observations, our DMD variants learn finite-rank operators from functional data such as observables, densities, or wavefunctions. We show that conventional DMD algorithms can be regarded as special cases of their functional DMD counterparts. All results will be illustrated with the aid of guiding examples. We focus in particular on Koopman, Perron-Frobenius, and Koopman-von Neumann operators associated with graphons, ordinary differential equations, and stochastic differential equations.

Thu 24 SeptMachine Learning
The gist
Dynamic mode decomposition (DMD) is a technique to simplify and understand complex changing systems by breaking them into patterns. The authors extend DMD to work directly with infinite-dimensional data, like continuous functions, rather than limiting it to simpler finite data. This lets them analyze systems modeled by equations involving continuous spaces without approximating them as finite sets. Their approach connects to important mathematical operators used to describe dynamics in areas like networks and random processes, helping to generalize DMD for broader scientific use.
Open → 2609.29159v1

Physical field method improves AI for predicting complex physics systems

PhyMo: A Physical-Field Modality for Multimodal AI4Physics

Abstract: Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.

Wed 23 SeptMachine LearningArtificial Intelligence
The gist
Predicting how physical systems behave often requires combining different types of data and knowledge. Existing AI methods treat physical measurements and equations as simple numbers or text, missing important physical rules that guide how these parts interact in space and time. This work introduces a new way to represent physical data called a physical-field modality, which respects these physical rules. The researchers built PhyMo, a system that first learns to understand physical fields, then aligns this knowledge with visual data, and finally uses the combined information to make better predictions. Tests on multiple physics problems show that PhyMo outperforms previous methods in understanding and predicting physical phenomena.
Open → 2609.27554v1

Time series forecasts improve by learning from past outcomes

When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting

Abstract: Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the forecasts it produces and the orchestration policy that determines which components to trust and how to coordinate them. The deployment process naturally provides supervision for this adaptation as forecast horizons elapse and realized targets reveal the effectiveness of earlier decisions. Committing all numerical expert forecasts and candidate agent paths before target observation allows each realized outcome to evaluate the entire alternative set, providing delayed feedback without additional annotation. However, existing time series agents primarily incorporate prior experience through forecast refinement, reflection, or retrieval, without systematically converting realized outcomes into persistent updates to the joint orchestration policy governing later origins. To exploit this delayed feedback systematically, we introduce TimEvolve, a frozen-backbone time series agent that converts each realized outcome into persistent joint updates of expert trust, agent path selection, and intervention strength. A temporally ordered predict, reveal, and update protocol applies this feedback to subsequent forecasts. Experiments across eight Time-MMD domains show that TimEvolve achieves the best average MSE and MAE ranks among fifteen methods and the lowest errors on both metrics in seven domains. These results demonstrate the value of learning forecasting policies from the futures encountered during deployment.

Mon 21 SeptMachine LearningArtificial Intelligence
The gist
Predicting future trends can be tricky when the system itself changes over time, making some methods better at certain moments. The authors introduce TimEvolve, a new forecasting system that learns from actual past results to improve future predictions and decides which forecasting methods to trust based on experience. Unlike earlier approaches, TimEvolve adapts its strategy continuously by using feedback from real outcomes to update how it combines models and interventions. Tests show it outperforms many existing methods across several scenarios where the underlying systems evolve.
Open → 2609.24862v1

Next generation reservoir computing infers missing parts of complex systems

Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data

Abstract: We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and Rössler system, where two unknown components are inferred from one given component. For both systems, NGRC achieves accurate results while requiring fewer training data and less computational time than RC. We identified an inverse proportional behavior between the number of time-delayed steps needed for NGRC and the temporal resolution, indicating that the physical time span covered by the delay interval is an important factor in determining the required number of delayed steps. Finally, we apply NGRC to the observational climate data of ENSO (El Niño--Southern Oscillation) and infer one observable from the remaining variables. Despite the noise and complexity of the real-world data, the NGRC shows promising results. Our findings demonstrate the potential of NGRC for efficient inference of unseen components in both controlled dynamical systems and real-world data.

Mon 21 SeptMachine Learning
The gist
Many systems, like the weather or climate, have parts we cannot directly observe, making prediction challenging. The authors used a method called next generation reservoir computing to guess missing parts of these systems from available data. They tested this on well-known chaotic systems and real climate data, showing it works well using less data and time than older methods. This approach might help better understand and predict complex natural phenomena by filling in unknown details.
Open → 2609.24754v1

Chronosphere adapts climate data detail across space and time

Chronosphere: Space-Time Tessellation of Local Climate Experts

Abstract: We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only predefined regions. Learned tessellations adapt, but are inefficient at representing higher frequencies. Chronosphere unifies these approaches, pairing an adaptive tessellation of learnable sites on the spacetime torus $S^2\times S^1$ with a shared bank of local basis functions. Both where capacity is placed and how much detail each region carries adapt to the data, across space and time. Trained to reconstruct climatology, Chronosphere matches or leads state-of-the-art location encoders across spatial and temporal tasks, with the largest gains under spatial and temporal transfer.

Fri 18 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Representing climate data is challenging because weather and climate change in complex ways across different places and times. The authors created Chronosphere, a machine learning method that divides space and time into adaptable regions, each capturing the right amount of climate detail. This approach combines local and global methods to better focus on important areas and times, improving how well climate information is understood and predicted. Their method outperforms existing ones, especially when applied to new locations or times.
Open → 2609.21872v1

OceanMoE improves long-term ocean variable forecasting with adaptive models

OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting

Abstract: Multivariate ocean forecasting must exploit shared evolution in a coupled ocean system while adapting to the heterogeneous statistical and dynamical characteristics of different prediction variables and locations. Fully shared models may lack the flexibility to handle this heterogeneity, whereas fully independent models discard the common ocean context shared across variables. The key question is how to retain shared context in a unified model while allowing computation to specialize according to the prediction target and local state. We propose OceanMoE, a structured conditional sparse Mixture-of-Experts framework that combines sharing and specialization for multivariate ocean forecasting. OceanMoE fuses cross-variable information to construct target-specific local representations and uses them to perform content-conditioned sparse routing at each spatial location, with the number of active experts adapted to router confidence. In the decoder, routing is augmented with a learned geographic bias parameterized by spherical-harmonic spatial bases, while shared residual and seasonal pathways provide common cross-variable and month-dependent context. Experiments on long-horizon autoregressive ORAS5 forecasting show that OceanMoE lowers aggregate forecasting error in both evaluated settings and maintains lower geometric-mean normalized RMSE than the corresponding baselines over most later rollout months. Routing analyses further show that expert allocation varies with prediction targets and spatial locations. These results support structured conditional computation as a modeling strategy for balancing shared ocean context with adaptive specialization.

Thu 17 SeptMachine Learning
The gist
Forecasting multiple ocean features like temperature and currents together is tricky because they change differently across locations and variables. The paper presents OceanMoE, a model that balances shared ocean information with specialized predictions by selecting the right small groups of experts for each place and variable. This approach helps make better long-term ocean predictions by adapting to local and variable differences while still learning from common ocean patterns. Tests showed OceanMoE reduced forecast errors compared to traditional methods and changed how experts were chosen based on what was being predicted and where.
Open → 2609.19768v1

Variational latent models improve long horizon pde predictions

Stable by Construction: Variational Latent Markov Operators for Long-Horizon PDE Prediction

Abstract: Neural PDE solvers provide efficient surrogates for time-dependent physical systems, but autoregressive prediction over long horizons remains challenging because local errors can induce distribution shift and accumulate under recursive deployment. We develop a variational approach to this problem by introducing latent Markov dynamics in which physical states are represented by latent distributions and evolved through probabilistic transitions. The framework is formulated directly on function spaces and specialized to functional Gaussian models, where structured latent perturbations induce a spectral geometry and variational transition alignment regularizes the learned dynamics. We further analyze how these mechanisms affect autoregressive error propagation, providing a theoretical connection between variational training and long-horizon prediction. We instantiate the framework as the Variational Autoencoding Markov Operator (VAMO), which combines spatially resolved latent fields, structured Gaussian perturbations, and a neural-operator transition. Empirically, we demonstrate the effectiveness of VAMO on several fluid-dynamics benchmarks with prediction horizons extending substantially beyond those represented during training, where it consistently reduces error accumulation and improves rollout stability over several deterministic and noise-injection baselines. Overall, these results highlight variational modeling as a complementary approach to robust long-horizon neural PDE dynamics.

Tue 15 SeptMachine Learning
The gist
Predicting how physical systems change over time using neural networks is hard because small errors add up. The authors developed a method that represents the system in a way that focuses on probabilities and smooth transitions, helping avoid growing mistakes. Their approach models the system’s hidden states with special math tools called variational latent distributions and Gaussian processes. This helps their predictions for things like fluid flow stay accurate even when looking far into the future.
Open → 2609.16621v1

Physics inputs speed up glacier flow simulations on single GPUs

Physics-enriched neural solvers for transient ice-flow simulation

Abstract: Transient glacier simulations with higher-order ice flow require the repeated solution of a nonlinear problem as the geometry evolves. In the online mode of the Instructed Glacier Model, the velocity field is represented by a neural network whose weights are warm-started from the previous time step and updated with a few optimizer iterations. We show that supplying the network with inexpensive input fields derived from low-order ice-flow balances improves this online solver. Unlike residual-based physics-informed neural networks, which incorporate physics through governing-equation penalties in the loss, our approach leaves the governing energy objective unchanged, adding physical structure through the network inputs. Across three real-world glacier configurations, the enriched solver is markedly more robust to solver settings. On the two alpine cases, it also improves the tuned accuracy--runtime trade-off, reducing surface-velocity errors by factors of two to four at fixed runtime and reaching few-percent relative errors with only $10^4$--$10^5$ trainable parameters, far fewer than comparable raw-input baselines. A 300-year Aletsch simulation then completes in under one minute, and the larger Valais domain in about two minutes, on a single GPU---a budget once reserved for much simpler shallow-ice models. Gains are smaller for the fast marine-terminating glacier, where nonlocal stress coupling favors larger or spectral networks. More broadly, the results suggest that enriching a neural solver's inputs with reduced-order physics can make repeated higher-order solves much cheaper, with no training data and no offline training.

Fri 11 SeptMachine Learning
The gist
Glacier flow simulations can be very slow because they require solving complex physics equations repeatedly as the glacier shape changes. The authors improved a neural network-based solver by giving it simple physics-based hints as inputs, which helps the network solve these equations faster and more accurately without needing extra training data. Their method runs fast enough to simulate hundreds of years of glacier movement in just minutes using a single GPU. This work shows a way to make advanced glacier models much more efficient for real-world use.
Open → 2609.12900v1

Bayesian neural network achieves faster precise weather forecasts with uncertainty

4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric Modeling

Abstract: We present BEAST, the first-ever Bayesian Swin Transformer for atmospheric forecasting on 0.25$^\circ$ global resolution able to accurately quantify both aleatoric and epistemic uncertainty. To overcome the associated computational bottlenecks, we devise an orthogonal 4D-parallelization scheme that introduces a unique domain-tensor-parallelism strategy and a novel uncertainty parallel method, enabling us to fully leverage GPU capacity and efficiently scale model training. For a 2.4-billion-parameter model, we achieve a peak performance of 3.96 EFLOP/s on 20,480 NVIDIA GH200 GPUs on the JUPITER supercomputer. We train BEAST as a 700-million-parameter model with 96 random weight samples on 384 nodes on 40 years of data for nearly one million gradient updates. This model achieves predictive skill scores competitive with state-of-the-art probabilistic atmospheric AI models and numerical models, and can predict extreme events with exceptional skill, while generating large ensembles 3 to 4 times faster than the current-best AI model. Our contribution unlocks the potential of high-fidelity uncertainty quantification in atmospheric AI models, heralding a new era for AI-based models in climate and Earth system sciences.

Fri 11 SeptArtificial IntelligenceDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Weather forecasting models need to predict not just what will happen, but also how sure they are about their predictions. The authors created BEAST, a new kind of AI model that uses a special neural network to make highly detailed weather forecasts across the globe and shows how confident it is in those forecasts. They also developed a way to run this huge model efficiently on many GPUs at once, making training faster and more powerful. Their results show the model can predict extreme weather well and generate many forecast possibilities more quickly than other AI models.
Open → 2609.12815v1