Model and dataset improve water level estimates in sparse river networks
A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph
Machine Learning
Summary
Measuring water levels in rivers is important for predicting floods and managing water but is hard because there are very few sensors on many rivers. The authors created a large dataset of satellite measurements over many rivers in the Amazon basin to help estimate water levels where data is missing. They found existing methods did not work well for this sparse and complex data, so they made a new model that uses river network structure and time together to make better guesses. Their model predicts water levels more accurately than previous approaches, especially in places without nearby satellite data.
What this means in practice
- •For water resource managers: Improve water level predictions for river sections lacking direct sensor data to better manage resources and anticipate floods.
- •For hydrological modelers: Use the new dataset and model to enhance hydrological simulations over large, sparsely measured river networks including the Amazon basin.
Authors
Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, Sébastien Lefèvre, Diego Fernandez Prieto
Abstract
Continuous monitoring of water surface elevation across river networks is critical for flood forecasting, water resource management, and understanding the global water cycle. Yet, the scarcity of in situ gauges across much of the globe constrains the development of reliable modeling frameworks. Satellite altimetry has the potential to alleviate this problem but its use is currently hindered by sparse temporal coverage. To this end, we introduce AmazonSWE, a dataset for training and evaluating large-scale spatiotemporal graph imputation methods that integrates processed satellite altimetry measurements from a range of sources, including the recent wide-swath SWOT sensor. The dataset covers over 19K river sections and 10 years (2016-2026) in the Amazon river basin, with in situ gauges held out for evaluation. Besides contributing a novel real-world use case with the potential for societal impact, AmazonSWE introduces significant technical challenges: with fewer than 1% of sections observed per day, the dataset is far sparser than existing imputation benchmarks, and its directed acyclic river topology is both structurally different from and larger than graphs in existing datasets. We show that prior spatiotemporal graph imputation methods are not adapted to this topology, scale and sparsity, and propose a simple bidirectional selective state space model that outperforms them by sampling connected subgraphs and flattening space and time into a single token sequence with topology-aware positional encodings. Compared to the state-of-the-art published method for SWOT-based WSE densification, which integrates statistics with physical modeling, our model reduces RMSE against in situ gauges by 18-39%, while producing predictions for every river section rather than only those with sufficient nearby satellite coverage.