Papers for

earth observation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Atomizer IO enables flexible processing beyond image grids

Atomizer-IO: Beyond Pixels, Patches and Grids

Abstract: Most vision architectures assume that observations lie on a regular grid, an effective abstraction for natural images but a restrictive one for sensing data whose channels, temporal sampling, spatial resolution, and geometry can vary. Generic set-based architectures remove the grid, but also remove useful spatial inductive biases. We introduce Atomizer-IO, an architecture that places observations first and derives structure from their physical relationships. Building on top of an atomic representation of the data, each observation is described by its measurement and acquisition metadata, while local cross-attention maps observations to anchor points that can be arbitrarily placed. We evaluate this design by progressively relaxing the grid assumption, from varying input raster configurations and incomplete channel sets to flexible output density and, ultimately, inputs without a raster grid. Atomizer-IO is competitive with flexible EO-specific architectures on most tasks, while offering post-training control over inference cost and competitive compute--performance trade-offs. The same formulation extends without architectural redesign to unordered 3D point clouds, showing that the atomic interface generalizes beyond regular raster inputs. These results suggest that pixels, patches, and grids do not need to define the interface of a sensing architecture.

Wed 30 SeptComputer Vision and Pattern Recognition
The gist
Most vision systems assume data comes in neat grids like pixels in a photo, which works well for pictures but not for many other types of sensor data that vary in shape or timing. The authors propose Atomizer-IO, a system that treats each data point as an individual unit with its own information and relates them through spatial connections rather than fixed grids. This approach works well even when input data changes formats or comes from unordered 3D point clouds, showing flexibility in handling various sensing setups. It performs comparably with specialized systems designed for specific earth observation tasks and allows control over computing resources during use.
Open → 2609.40320v1

Planetary feature fields compress earth data with speed and accuracy

Planetary Feature Fields are Scalable Earth Representations

Abstract: Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume---a decomposition of an explicit 3D grid with smaller factors---across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At $1800\times$ compression relative to the uncompressed source data, reconstructed features retain approximately $90\%$ or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.

Tue 29 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Satellite data about Earth comes from many sources and is very large, making it hard to store and use together. The authors present Planetary Feature Fields (PFFs), a new way to combine these data into a compact, continuous representation that keeps important detail over time and space. PFFs compress data by over a thousand times while still performing well on tasks like recognizing land features or changes. They also allow adding new data easily and let users access information much faster than existing methods.
Open → 2609.37784v1

Hypersam builds general model for analyzing earth images

HyperSAM: A Promptable Foundation Model for Hyperspectral Remote Sensing

Abstract: Hyperspectral remote sensing provides dense spectral measurements that are indispensable for material-level Earth observation, yet the construction of a general-purpose hyperspectral foundation model remains difficult. Two bottlenecks are especially limiting. First, large hyperspectral corpora rarely provide high spatial resolution together with reliable dense annotations. Second, many hyperspectral models are still trained almost from scratch, so the geometric and interactive priors learned by modern vision foundation models are not fully reused. To alleviate these issues, we \highlight{present} \textbf{HyperSAM}, a promptable hyperspectral foundation model that couples a data-centric hyperspectral synthesis pipeline with a spectral adaptation architecture based on Segment Anything Model 3 (SAM3). On the data side, HyperSAM synthesizes full-spectrum hyperspectral cubes from high-resolution SpaceNet multispectral imagery through a physics-informed abundance-transfer generator, while SAM3-derived pseudo-masks provide object-centric supervision. On the model side, the latest implementation uses a frozen SAM3 RGB image branch, a trainable hyperspectral side encoder initialized from the RGB vision transformer (ViT), ControlNet-style zero-initialized feature injection, and a lightweight mixture-of-experts mask refiner. To enhance training robustness against noisy pseudo-labels, Cross-modal Sample Selection (CromSS)-style confidence selection is incorporated for noisy-label weighting. Extensive experiments show that HyperSAM obtains strong generalization on diverse hyperspectral tasks (e.g., classification, anomaly detection, change detection, target detection, and airborne oil-spill mapping) and that high-quality synthetic hyperspectral data can be more effective than simply scaling noisy hyperspectral supervision.

Tue 29 SeptComputer Vision and Pattern Recognition
The gist
Hyperspectral remote sensing collects detailed color information from the earth's surface, which helps identify materials and track changes. The authors created HyperSAM, a new computer model that can work with these complex images more effectively by combining real data and synthetic images. It uses another model called SAM3 to better understand objects in the images and smart techniques to handle imperfect labels. This approach lets HyperSAM perform well on various tasks like detecting changes, finding targets, and spotting anomalies in earth observation data.
Open → 2609.37340v1

Earth observation models change more when fine-tuned than natural image models

Reuse or Relearn? A Spectral View of Earth Observation Foundation Models

Abstract: Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reuse the pretrained representation, or does it relearn a new one? We study this with spectral diagnostics that compare a model before and after adaptation, quantifying how well its dominant singular subspaces are preserved, how broadly the weight update is distributed, and how large it is. Using natural image models such as CLIP and DINO as a reference, we find that, under the evaluated fine-tuning settings, EO models undergo far larger, higher-rank updates and retain much less of their pretrained structure, so their downstream performance is often obtained with substantial changes to the pretrained weight structure. The diagnostics further provide insight into how cheaply a model can be adapted: where the pretrained subspaces are preserved, adapting a small fraction of the parameters can match full fine-tuning, and where they are not, it can fall behind. More broadly, foundation models, and EO foundation models in particular, should be assessed not only by benchmark accuracy, but also by how reusable their pretrained representation is.

Sat 26 SeptMachine Learning
The gist
Fine-tuning is a common method to adapt big AI models to specific tasks. The authors studied if fine-tuning changes Earth observation (EO) AI models a lot or keeps their original knowledge. They found EO models change more extensively compared to models trained on everyday photos. This means EO models often learn new information rather than just reusing what they already knew. Their findings suggest it’s important to evaluate these models not just by how accurate they are, but also by how much of their original knowledge is preserved.
Open → 2609.32756v1