GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation

2026-08-03Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors created GeoCore-9B, a large model trained only on earth observation (EO) data, unlike earlier models that used general images and had biases. Their model uses a special method called a Flow Matching-based Diffusion Transformer and includes location details like latitude and longitude to generate more accurate images. They introduced a new training technique called Geospatial Semantic Alignment loss to help the model understand Earth-specific features better. GeoCore-9B performs well on tasks like removing clouds from images and converting radar images to optical ones, showing strong results in both image quality and geographic accuracy.

earth observation (EO)generative modelFlow Matching-based Diffusion TransformerGeospatial Semantic Alignment lossgeospatial metadatadiffusion modelcloud removalSAR-to-optical translationfoundation modelGit-10M dataset
Authors
Jeonghyeok Do, Munchurl Kim
Abstract
Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes. To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead. Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy.