Supraglacial Lake Fate Is Knowable Long Before the Season Ends

2026-08-31Machine Learning

Machine Learning
AI summary

The authors studied how supraglacial lakes on the Greenland Ice Sheet end their melt seasons in four ways: rapid drainage, slow drainage, refreezing, or burial by snow. They found that two outcomes (rapid and slow drainage) can be predicted much earlier in the season—about two to three months before the full data is available—using satellite data. This means monitoring systems don’t have to wait until the end of the melt season to know important details about lake outcomes. Their results were consistent across different machine learning models and areas. However, predictions were less reliable when using automatic labels instead of expert ones for unfamiliar seasons.

supraglacial lakeGreenland Ice Sheetmelt seasonrapid drainageslow drainagesatellite classificationmachine learninghydrofractureearly predictionrefreezing
Authors
Emam Hossain, Md Osman Gani
Abstract
A supraglacial lake on the Greenland Ice Sheet ends its melt season in one of four ways: it drains rapidly through a hydrofracture, drains slowly across the surface, refreezes in place, or is buried by late-season snowfall. Which one occurs decides whether the meltwater reaches the ice bed. Satellite classifiers recover the outcome accurately but only after the season closes, and how much of a season each outcome actually requires has never been measured. We measure it directly: holding the representation and the classifier fixed, we truncate the input at $14$ cutoffs from 1 May to 31 December, retrain at each, and record the earliest cutoff at which each outcome's per-class $F_1$ reaches a fixed target. The outcomes resolve in a consistent order, two of them months early: rapid drainage by 15 July and slow drainage by 1 August, $92$ and $75$ days ahead of the earliest date a full-season pipeline can be computed at all, with buried and refreeze following at $44$ and $30$ days. Five further learners, from a majority-class floor and $54$ summary statistics to a trigger-based early classifier, leave the ordering intact: every learner that produces a per-class trajectory reproduces it despite end-of-season accuracies differing by up to $18$ percentage points, and it survives leave-one-basin-out evaluation, though not the substitution of machine labels for expert ones in an unseen season. Every feature we compute at day $t$ reads only days up to $t$, at a cost of at most $1.3$ percentage points. A monitoring system should therefore not have one release date: rapid drainage can be flagged on 15 July, three months before a full-season pipeline can be computed at all.