Papers for

weather forecasting teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Extreme event prediction improves with local instability sensing

Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

Abstract: Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only \(50\) time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with $20$ times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.

Fri 25 SeptMachine Learning
The gist
Predicting rare, extreme events in chaotic systems is tough because these events depend on short-term instabilities that are not common in usual data. The authors developed a method that uses a group of slightly different simulations to detect these local instabilities without complex calculations. By feeding this information into machine learning models, the method improves the prediction of rare events, even with limited training data. They tested their approach on both simple chaotic systems and more complex fluid models, showing better accuracy compared to traditional methods.
Open → 2609.30746v1

WeatherDiagFlow improves radar forecasts with evidence and audits

WeatherDiagFlow: Evidence-Grounded Radar Nowcasting with Diagnostic Flow Refinement

Abstract: Radar nowcasting is essential for short-term warning and emergency response, yet conventional systems mainly return future radar fields and provide limited support for operational communication and post-event verification. We formulate radar nowcasting as an evidence-grounded forecast--bulletin--audit task, in which a numerical forecaster produces both future radar fields and structured diagnostic evidence. Forecast-time bulletins use only model-available evidence, whereas post-event audits incorporate future radar truth only after the forecast horizon is observed. Based on this task formulation, WeatherDiagFlow predicts motion, growth and decay, heavy-echo risk, and uncertainty to condition rolling flow refinement, while frozen-scaffold residual calibration improves long-lead strong-echo preservation. A multi-agent layer converts the structured evidence into operational bulletins and independently generates verification audits without feeding textual outputs back into the forecaster. Experiments on FJRADAR demonstrate competitive overall performance and improved strong-echo event skill. WeatherDiagFlow therefore connects numerical prediction, evidence-grounded reporting, and auditable verification under a leakage-controlled protocol.

Thu 24 SeptMachine LearningMultimedia
The gist
Short-term weather forecasts from radar help with emergency warnings but usually only show future radar images without much explanation. The authors present WeatherDiagFlow, which not only predicts future radar but also provides clear evidence and reports about what the forecast means and how confident it is. This system also checks its own forecasts against real radar data afterward to verify accuracy. By combining predictions, evidence, and audits, WeatherDiagFlow aims to make radar forecasts easier to understand and trust.
Open → 2609.29772v1

Wavelet-diffusion models improve precipitation detail across US regions

Evaluating Cross-region Generalization for Wavelet-Diffusion Precipitation Downscaling

Abstract: Diffusion models have shown strong potential for kilometer-scale precipitation downscaling, but their performance in geographically unseen regions and event regimes remains insufficiently understood. Building on the wavelet diffusion model (WDM) framework, this study evaluates cross-region and cross-event generalization. Six 3 x 3 deg U.S. regions represent convective, winter, tropical, and atmospheric-river precipitation regimes. Low-resolution inputs are generated by block averaging NOAA Multi-Radar/Multi-Sensor (MRMS) composite reflectivity fields. A WDM trained only on Oklahoma (OK) samples and a WDM trained on all six regions are compared with nearest-neighbor and Bicubic interpolation. Model performance is evaluated using three metric families that measure image-domain reconstruction, spectral and distributional fidelity, and bin-wise precipitation detection. The OK-trained WDM remains competitive outside OK. Although the all-region WDM delivers the best and most consistent overall image-domain and detection performance, its gains are uneven across precipitation intensities. Bin-wise critical success index (CSI) over 5-dBZ reflectivity bins shows that WDM improvements concentrate in localized higher-reflectivity structures, which image-domain metrics partly obscure. In addition, the performance differences among samples are strongly associated with the spatial organization of the precipitation field, quantified by Moran's I as the spatial autocorrelation of each reflectivity bin. The sample-level Moran's I-CSI correlation stratified by sample intensity reaches 0.901 in all six regions, including regions unseen during training. Overall, these findings support future efforts to transfer downscaling models to regions with limited local training data and to generate globally consistent, high-resolution precipitation products.

Wed 23 SeptMachine Learning
The gist
Predicting detailed rainfall patterns from coarse weather data is hard, especially in places where models weren’t trained. This paper looks at a special AI method called wavelet diffusion models (WDMs) that can create detailed precipitation maps from smooth inputs. The authors test how well these models work in different US climate regions, including places they haven’t seen during training. They find that WDMs trained on one region can still do a good job elsewhere, and models trained on all regions work even better in spotting intense rain areas. This could help make weather predictions more detailed and accurate in many places, even with limited local data.
Open → 2609.28749v1

Transformer model improves planetary boundary layer height estimates from satellite data

PBLH Estimation from Satellite Radiances via a Dual-Encoder Transformer

Abstract: Estimating the Planetary Boundary Layer Height (PBLH) from satellite observations is a challenging regression problem due to the indirect relationship between top-of-atmosphere radiances and near-surface atmospheric structure. Progress has been limited both by the lack of architectures capable of handling the multimodal, spatially incomplete nature of satellite overpasses, and by the scarcity of suitable datasets. In this paper, we build upon the large-scale dataset pairing MetOp radiances with ERA5 PBLH labels that we introduced in our previous work, making three contributions. First, we establish a benchmark across eight approaches spanning pixel-wise regression, swath-wise sequence models, and convolutional and Transformer models operating on the full orbital passage. Second, we quantify what the resulting model actually relies on, using grouped Shapley decomposition over the input blocks. Third, we present the best-performing architecture found: a dual-encoder Transformer whose masked-input handling lets it operate in all weather conditions. The proposed model achieves MAE = 155.8 m on the held-out global test set, outperforming all baselines on every evaluation subset. On 30 out-of-distribution granules acquired on two days overlapping the TEAMx observational campaign, it achieves MAE = 165.3 m, outperforming a pixel-wise baseline trained on the same data (MAE = 197 m).

Wed 23 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Estimating how high the lowest part of the atmosphere reaches (called the boundary layer) using satellite data is tricky because satellites see indirect signals. The authors worked with a big dataset matching satellite readings to known boundary heights and tested multiple methods to find the best one. They designed a special artificial intelligence model that looks at the data in two ways and can handle missing or unclear satellite readings caused by clouds or weather. Their model made more accurate height predictions worldwide than previous methods, including on new data taken during a recent weather campaign.
Open → 2609.28286v1

Sea ice forecast errors corrected using adaptive sparse observations

Taking a Second Look: Correcting Sea Ice Forecasts with Sparse Observations

Abstract: Sea ice forecasts are issued several days ahead, allowing errors to accumulate while new, often sparse sea ice concentration (SIC) observations become available. We find that fixed-propagation errors concentrate near structured, high-gradient ice edges, whereas homogeneous interiors require limited propagation, suggesting that propagation distance should be state dependent. We therefore introduce ECHO (Evidence-guided Correction with Heterogeneous prOpagation), where ECHO-Scale adapts propagation distance while preserving correction geometry, and ECHO-Delta learns a bounded residual around fixed propagation. Across all 96 standard evaluation settings spanning diverse priors, observation times, sparsity levels, geometries, and noise conditions, both outperform fixed propagation. ECHO-Delta achieves the best average accuracy, while ECHO-Scale is more robust to geometry shifts. Code is available at https://github.com/yingtian22/TAKING-A-SECOND-LOOK.

Mon 21 SeptMachine Learning
The gist
Forecasts of sea ice several days ahead can have errors, especially near ice edges where conditions change quickly. The authors found that these errors spread differently across ice areas, depending on whether the ice is stable or near borders. They created ECHO, a method that adjusts how far correction information travels based on ice conditions, improving forecast accuracy. Their approach works better than traditional methods across many different scenarios, handling sparse and noisy data more effectively.
Open → 2609.24591v1

Recursive quantum lstm improves temperature prediction accuracy and stability

Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting

Abstract: Quantum long short-term memory (QLSTM) models extend recurrent sequence learning with variational quantum circuits, but their optimization behavior can vary substantially across random initializations and temporal contexts. This paper evaluates a recursive QLSTM architecture against a standard QLSTM for one-step-ahead prediction of daily minimum and maximum temperature. Using daily weather observations from Toronto and identical training settings, we compare convergence, predictive accuracy, and generalization across input windows of 8, 16, and 32 days over 20 random seeds. The recursive model consistently reaches a near-optimal test loss earlier, reduces mean absolute error and root mean squared error, and exhibits a smaller generalization gap. These results indicate that recursive quantum feature transformations can improve stability and out-of-sample performance for compact hybrid quantum--classical temporal models.

Thu 17 SeptMachine Learning
The gist
Predicting daily temperatures can be tricky, especially when using advanced computer models that mix quantum computing ideas with traditional methods. The authors studied two versions of these models to see which predicts daily highs and lows better for Toronto's weather. They found that a recursive version, which reuses information in a special way, was more stable and made more accurate forecasts than the regular version. This suggests new ways to improve weather predictions using hybrid quantum-classical techniques.
Open → 2609.20594v1

Double U-shaped neural operator improves wave modeling accuracy and efficiency

DU-NO: A Parameter-Efficient Double U-Shaped Neural Operator for Phase-Resolving Wave Modeling

Abstract: Phase-resolving wave models such as FUNWAVE-TVD are the accuracy standard for nearshore dynamics, resolving the shoaling, refraction, and breaking of individual waves, but their cost rules them out for the ensembles, uncertainty quantification, and real-time warning that operational forecasting demands. Neural operators promise solver-level accuracy at a fraction of that cost, yet on wave-dominated fields the accurate ones are large: hybrid spectral-convolutional operators such as U-FNO (the strongest baseline in our study after DU-NO) buy their fidelity with tens of millions of parameters. We introduce DU-NO (Double U-shaped Neural Operator), a multiscale U-shaped spectral operator that attaches lightweight convolutional U-Net branches only at its two shallowest encoder and decoder levels. The placement follows a sampling argument: high-wavenumber content exists only on fine grids, so the local, full-band pathways go where that content lives, while the coarse, band-limited levels stay purely spectral. A depth-decaying mode schedule holds the model to 3.64M parameters, an order of magnitude below U-FNO. On our publicly released FUNWAVE-TVD benchmark, DU-NO attains the best autoregressive rollout error of six identically trained architectures, improving on U-FNO by 14.9% with 10.8x fewer parameters, and a frequency-band analysis shows the gain holds across all bands, including the high-wavenumber band where truncated-spectral operators collapse. Parameter-matched controls confirm the gain is architectural: rescaled to the same 3.6M budget, the best baseline still trails DU-NO by 28.6%. The advantage carries beyond nearshore waves: DU-NO matches the strongest baselines on 2D Navier-Stokes and wins clearly on PDEBench shallow-water rollouts. Code, trained models, and evaluation artifacts are available at https://anonymous.4open.science/r/duno-code-5A7B/.

Thu 10 SeptArtificial Intelligence
The gist
Modeling how waves behave near shores can be very accurate but usually requires slow and expensive computer simulations. The authors developed a new neural network design called DU-NO that is much smaller and faster but still very accurate at predicting waves. DU-NO uses a clever structure that focuses on fine details where needed, reducing the number of parameters by about ten times compared to strong previous methods. It also performs well on other fluid dynamics problems, suggesting broad usefulness.
Open → 2609.12115v1

Post-processing methods improve temperature and wind forecasts at new locations

Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts

Abstract: Statistical post-processing improves ensemble weather forecasts, but generating calibrated predictions at locations without observations remains challenging. This study compares statistical and machine-learning-based methods for post-processing ECMWF 2-m temperature and 10-m wind speed forecasts at observed and unobserved stations in Germany. We consider EMOS-based approaches, distributional regression networks, Transformers, and graph neural networks under both limited and extended predictor settings. For temperature, we also investigate linear forecast combinations and propose an altitude-aware linear pool (ALP). The results show that post-processing improves upon the raw ensemble in most settings, but no single method performs best across all variables, station groups, and evaluation metrics. The proposed ALP provides a small but significant improvement over the standard linear pool at unobserved locations.

Mon 7 SeptMachine Learning
The gist
Weather forecasts often use groups of predictions that need adjusting to be more accurate. The authors compared traditional statistical methods to newer machine learning techniques to improve forecasts of temperature and wind speed in Germany, especially at places without weather stations. They found that while no single approach worked best for every situation, adjusting forecasts this way generally made them better than raw predictions. The authors also introduced a new way to combine forecasts that takes altitude into account, which slightly improved temperature predictions at unobserved locations.
Open → 2609.07512v1