Papers for

remote sensing teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Spectral super-resolution improves satellite image detail and color accuracy

Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks

Abstract: Spectral super-resolution of multispectral satellite images can enable high temporal- and spatial-resolution hyperspectral satellite imagery at a modest cost, significantly increasing the applicability of hyperspectral remote sensing. This task is inherently ill-posed, making it well-suited for deep learning-based methods. In this study, the spectral super-resolution task is framed as an operator learning problem, and SSRON is proposed as a Deep Operator Network that effectively learns function-to-function mappings from downsampled spectra to continuous spectra. The model is trained to super-resolve Sentinel-2A-like multispectral imagery to EMIT images. Compared to baseline models, SSRON achieves superior performance across all metrics. The model also demonstrates zero-shot spectral super-resolution capability by predicting bands unseen during training. Furthermore, its continuous-output formulation suggests the potential to estimate spectra at finer wavelength intervals than the native sensor. These results suggest the potential of SSRON and establishes operator learning as a promising direction for spectral super-resolution.

Mon 28 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Satellite images often capture limited colors or wavelengths, making it hard to see detailed information about the land or environment. The authors present a new deep learning method called SSRON that can take these simpler images and predict more detailed color information, similar to higher-quality hyperspectral images. This method is especially good because it learns continuous spectral information and can guess details even for colors it hasn't seen before. This could help scientists and businesses get better insights from satellite imagery without needing expensive sensors.
Open → 2609.35410v1

Preference guidance improves open-vocabulary image segmentation in new domains

Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement

Abstract: Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences. Our code is available at https://github.com/blue-531/pref-ovss.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Semantic segmentation assigns labels to every pixel in an image, but models often struggle when applied to specialized areas like medical images or satellite photos. This paper shows how asking simple yes-or-no questions about uncertain parts of an image can help improve these models without needing detailed, hard-to-get labels. The authors use disagreements between different text prompts as clues to figure out where to ask these questions, leading to better segmentation results. Their method works well on various existing models and remains effective even if some answers are noisy.
Open → 2609.34528v1

GeoCR removes clouds from satellite images across sensors and bands

GeoCR: Learning a Generalist Cloud Removal Prior from Heterogeneous Observations

Abstract: Cloud removal methods are typically specialized to individual datasets and input configurations, limiting reuse across sensors, spectral bands, and observation settings. We introduce GeoCR, a generalist model that unifies RGB-only-based CR and multispectral-based CR from single- or multi-temporal cloudy observations, with optional SAR guidance, within a single network. To accommodate different spectral and sensing domains, compact input and output stems extend a pretrained RGB autoencoder while keeping its encoder and decoder trunks frozen. This shared latent interface enables a single flow transformer to jointly model clean RGB and non-RGB latents, conditioned on separate cloudy-observation streams and optional SAR tokens. Through joint pretraining on the training splits of ten datasets comprising 883,331 cloud-free target images, GeoCR learns a shared cloud removal prior across these heterogeneous configurations. The same pretrained checkpoint supports direct inference without dataset-specific fine-tuning and efficient adaptation through low-rank adaptation (LoRA). We evaluate GeoCR against general image restoration and cloud removal methods on test splits of the contributing datasets under full-band and RGB-only settings. GeoCR achieves the best FID and DISTS on full-band SEN12MS-CR and Sen2_MTC_New and RGB-only CUHK-CR2, outperforming existing models and demonstrating the effectiveness of a reusable generative model across diverse settings.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Clouds often block the view in satellite images, making it hard to see the land below. The authors created GeoCR, a single model that can clear clouds from many types of satellite images, including those with different colors and radar data. GeoCR learns from a huge variety of cloud-free images so it can work well on new images without extra training. This means fewer clouded images and clearer pictures for tasks like monitoring the environment or mapping.
Open → 2609.32510v1

Hyperspectral video compression improves quality and tracking accuracy

Implicit Neural Representation for Hyperspectral Video Compression

Abstract: With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bjøntegaard Delta PSNR gains of +4.99 dB and Bjøntegaard Delta rate of -88.88% compared to traditional hyperspectral image compression methods applied frame-by-frame. In addition to reconstruction quality, the effects on downstream task performance are measured in the form of object tracking success. Compared to video compressed with methods based on principal component analysis and JPEG2000 in low data regimes, our proposed method improves tracking area under the curve by up to 23.42% and distance precision by up to 35.56% on examples from the HOT2026 dataset.

Fri 25 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
The gist
Hyperspectral videos capture more colors than regular videos but create huge data files that are hard to store and share. The authors improved how these videos are compressed using a special kind of neural network, making the files much smaller while keeping the video quality better than older methods. This also helped computer programs track objects in the videos more accurately. Their method works better, especially when there isn’t much data to work with.
Open → 2609.31435v1

Flow matching improves speed and detail in SAR to optical image conversion

ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning

Abstract: In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image translation. Unlike conventional diffusion models, ContraFM-S2O learns to predict the velocity field in training and solves ODE instead of SDE during inference to improve the sampling efficiency. In addition, ContraFM-S2O replaces instantaneous velocity with average velocity along the interpolation path to realize one-step SAR-to-optical image translation and uses contrastive learning to improve the quality of the generated optical images. Experiments show our model achieves state-of-the-art on SAR2Opt and QXS datasets, outperforming baselines, and reduces inference latency via one-step generation.

Fri 25 SeptComputer Vision and Pattern Recognition
The gist
Turning radar images (SAR) into regular photos that people can understand is hard because details often get blurry or lost in the process. The authors created a new method called ContraFM-S2O that uses a mathematical technique called flow matching to convert images faster and with clearer details. Instead of taking many slow steps, their approach predicts the full transformation in one go. They also use a learning method called contrastive learning to make the converted images more accurate and sharp. Tests show their method works better and faster than older methods on common datasets.
Open → 2609.31378v1

Selective tool use improves change question answering in satellite images

Selective Tool Use for Agentic Change Visual Question Answering in Remote Sensing

Abstract: Change visual question answering (Change VQA) requires understanding semantic changes across bi-temporal remote sensing images. Although vision language models (VLMs) have shown promising performance on this task, they remain unreliable when answering questions that require explicit transition statistics, area measurements, or spatial information. To address this limitation, we propose a selective tool use framework in which a single VLM either answers directly or invokes a deterministic change analysis tool to obtain question specific evidence. Specifically, the selected tool operates on bi-temporal semantic maps and returns a structured observation, which the same VLM uses to generate its final answer. To support this framework, we construct a tool augmented extension of CDVQA covering eight question families and three tools for transition, spatial, and temporal analysis. Tool use supervision and observations are derived automatically from the original semantic annotations, without additional manual labeling. We then adapt Qwen3.5-4B using Low Rank Adaptation (LoRA) to jointly learn direct answering, tool invocation, and evidence conditioned answering. Experiments on 7,164 test questions show that selective tool use with reference semantic maps improves overall accuracy from 73.77% to 88.79% and average family accuracy from 69.11% to 89.65%. When the semantic maps are predicted automatically, the framework achieves 77.47% overall accuracy and 75.06% average family accuracy. These results demonstrate the benefit of question-specific semantic evidence for Change VQA, while highlighting the influence of semantic prediction quality on the resulting performance. Code and tool-augmented annotations will be made publicly available at https://github.com/yakoubbazi/ToolChangeVQA.

Sun 13 SeptComputer Vision and Pattern Recognition
The gist
Answering questions about how land changes over time using satellite images can be tricky, especially when precise details like area size or location matter. The paper presents a method where an AI first decides if it can answer directly or if it should use special analysis tools that measure changes in the images. These tools provide clear evidence which helps the AI give better answers. This approach significantly boosts accuracy, especially when the analysis maps are precise.
Open → 2609.14523v1