Sea-ice type mapping uncertainty linked to expert disagreement and model confidence

Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts

Machine Learning

Summary

Different experts often disagree when labeling sea ice types because the observations they use can be interpreted in various ways. This makes training computer models to recognize ice types harder. The authors studied how uncertainty from experts' disagreement and uncertainty from the model itself relate to each other. They found that using data showing multiple experts' opinions helps the model better reflect where ice types are more uncertain, especially near the ice edge. Also, one method called Monte Carlo dropout gave the most reliable confidence scores.

What this means in practice

  • For marine navigation teams: Improve sea-ice maps with uncertainty estimates that highlight where ice conditions are less certain, helping safer route planning near the ice edge.
  • For environmental monitoring agencies: Use multi-expert labeled data to train automated systems that better capture disagreement in ice type assessments for operational ice condition monitoring.

Authors

Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett, Farnoush Banaei-Kashani

Abstract

Sea-ice stage of development (SoD) describes the age and associated thickness of sea ice and provides important information for navigation, and operational ice monitoring. SoD labels are obtained from operational ice charts, where trained analysts interpret satellite observations and assign standardized stage codes to regions with similar ice conditions. These codes often represent ranges of compatible ice thicknesses rather than exact physical values. Deep-learning methods can automate SoD mapping and commonly adopt operational ice charts as reference labels for training. These annotations are not exact, however; this is because chart interpretation relies on analyst judgement and on the observations available at the time, so different ice services may assign different SoD labels to the same conditions. We term this variation across independently produced expert annotations multi-annotator label uncertainty; collapsing the annotations into a single deterministic target discards this variation. A second source of uncertainty originates in the learned model itself. In this paper, we quantify both sources: annotation uncertainty from disagreement among independent ice-service charts and model uncertainty from the learned predictive models. We then evaluate their relationship by testing whether model uncertainty is higher where ice services disagree. We observe that supervision incorporating information from multiple annotators can improve this correspondence, with soft supervision achieving the highest overall correlation of 0.256. The relationship becomes substantially stronger near the ice edge, where model predictive uncertainty closely tracks multi-annotator disagreement, reaching a correlation of 0.704 within 0--10 km. Among the uncertainty-estimation approaches, Monte Carlo dropout provides the best-calibrated confidence estimates, with an expected calibration error of 0.050.