From Machine Learning to Large-Scale EO Products: Best Practices for Making Maps

2026-07-27Machine Learning

Machine Learning
AI summary

The authors explain how creating large maps using satellite data and machine learning is now easier, but still tricky. They highlight that decisions made early on, like how data is prepared or chosen, affect the final map’s accuracy. Their work describes best practices across the entire process, from handling satellite data to checking map quality and sharing results. The goal is to help others make better, more reliable Earth observation maps.

Earth observationmachine learninggeospatial mapsdata preprocessingdataset constructionuncertainty quantificationmodel trainingmap validationsatellite dataglobal-scale inference
Authors
Ghjulia Sialelli, Robin Young, Yuchang Jiang, Cesar Aybar, Linus Scheibenreif, Damien Robert, Clemens Mosig, Adam J. Stewart, Jan D. Wegner, Aleksis Pirinen, Olof Mogren, Konrad Schindler
Abstract
Recent years have seen a rapid expansion in the production of large-scale geospatial maps derived from Earth observation (EO) data, driven largely by advances in machine learning (ML) and large computing infrastructure. Although the barrier to generating such maps has dropped substantially, established best practices have yet to emerge, and design decisions made early in the pipeline can quietly propagate errors into the final product. Producing a technically sound and scientifically credible product remains challenging. Choices made at every stage are tightly coupled: preprocessing decisions shape the training signal, dataset design governs what the model can learn and how reliably its performance can be assessed, and global-scale inference introduces engineering challenges in compute and data access at scale, as well as artifact mitigation. Furthermore, uncertainty quantification and independent map validation each require dedicated methodological attention that is often underestimated. This paper presents a concise, end-to-end account of the recommended practices spanning the pipeline from satellite data to an operational map product. We organize the discussion around six interconnected themes: the EO data infrastructure landscape, data selection and preprocessing, ML dataset construction and model training, uncertainty quantification, map production and distribution, and validation. This paper is a condensed version of a longer guide that provides greater depth across all stages, accessible online at ghjuliasialelli.github.io/MLEO-Maps/.