Papers for

environmental monitoring teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Comprehensive ais dataset reveals maritime activity in finnish waters

A Large-Scale AIS Dataset from Finnish Water

Abstract: This research paper contributes to the maritime research community by introducing a comprehensive AIS dataset from Finnish waters, specifically the Baltic Sea region. AIS data, initially designed for collision prevention, have evolved into a versatile tool with applications across diverse maritime domains. Our paper not only curates and categorises existing AIS datasets but also introduces a collected AIS dataset from the Baltic Sea area, renowned for its intercontinental cargo routes, military activities, and frozen water expanses. This dataset includes 229 millions data points and provides researchers with a resource for studying maritime activities and vessel behaviour in this dynamic region, notably distinguished by the inclusion of data from Finnish lakes. Our analysis includes a detailed listing of ship types and relevant features, empowering researchers to explore various maritime domains. To enhance comprehension and analysis, we provide visualisations of maritime traffic patterns. By publishing this AIS dataset, we aim to catalyse innovation and collaboration in maritime research, offering a gateway to deeper insights into maritime activities in the Baltic Sea and Finnish lakes.

Fri 11 SeptMachine Learning
The gist
Tracking ship positions helps keep waters safe, but these signals can also tell us a lot about how ships move and behave. The authors collected a very large set of such tracking data from Finland’s lakes and the Baltic Sea, a busy and complex water area. This dataset includes information about different types of ships and where they travel, helping people study maritime traffic. They also provide maps showing common routes, making it easier to spot patterns and changes in ship movements.
Open 2609.12938v1

AquaCubeAI enables faster onboard monitoring of coastal water turbidity

AquaCubeAI-Powered Monitoring Turbidity on-board Φsat-2

Abstract: Timely monitoring of coastal water quality is critical for environmental protection, yet conventional satellite workflows rely on downlink and ground processing, introducing latency that can limit responsiveness to rapidly evolving turbidity events. To address this limitation, we propose AquaCubeAI, a lightweight machine-learning approach for onboard estimation of coastal water turbidity from Φsat-2 multispectral imagery. By shifting inference from the ground segment to the satellite, AquaCubeAI aims to enable lower-latency, more responsive, and more operationally useful turbidity monitoring under the strict compute and bandwidth constraints of spaceborne platforms. The model is trained on simulated Φsat-2 acquisitions spatially aligned with Copernicus Marine Service (CMEMS) High-Resolution Ocean Color (HR-OC) turbidity products over selected localized coastal sites spanning four European marine macro-regions. To provide a realistic evaluation of generalization in the presence of spatial correlation, we adopt a spatial block splitting protocol that mitigates data leakage between training and evaluation subsets. The main contributions of this work are: (i) a scalable dataset generation pipeline pairing simulated Φsat-2 multispectral patches with CMEMS HR-OC turbidity labels across selected localized European coastal sites; (ii) a compact Multi-Layer Perceptron (MLP)-based turbidity regressor trained under a leakage-aware geospatial split and tailored to embedded constraints; and (iii) a reformulation for dense spatial prediction via parameter sharing, enabling turbidity mapping and simple threshold-based anomaly masks for onboard decision logic. Embedded deployment on an Intel Myriad Vision Processing Unit (VPU) further confirms the feasibility of low-power hardware and supports low-latency inference from multispectral inputs.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Monitoring how clear or muddy coastal waters are is important for protecting the environment. Existing satellite methods take time because they send images back to Earth for processing. The authors developed AquaCubeAI, a small and efficient AI model that runs directly on a satellite to measure water turbidity faster. They trained and tested the model using simulated satellite images matched with real water quality data from European coastal areas. This approach can provide quicker alerts about changes in water quality using low-power hardware on satellites.
Open 2609.12744v1

Earth observation agents plan and execute workflows with improved accuracy

Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents

Abstract: Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare data, perform domain computations, and derive conclusions from runtime evidence. Existing EO agents typically start from supplied observations, while benchmarks typically provide prepared inputs or candidate answers, leaving full-chain open-world EO execution largely untested. We present Earth-Agent-Pro, an execution-adaptive Plan-and-Execute framework using expert-authored skills to constrain planning and runtime tool use. Workflow-centered structured memory records planned steps, accepted evidence, and their dependencies, enabling repair of only the affected workflow suffix when runtime evidence invalidates a step. Separate large language model adapters use sequence-level supervised fine-tuning for planner workflow composition and node-level group relative policy optimization with locally verifiable rewards for executor tool-argument grounding. Earth-Bench-Pro instantiates 248 expert-curated task cores as 744 questions under three matched regimes. Its 248 Open-World Execution questions span RGB imagery, spectral observations, and remote sensing products, pairing high-level requests with runtime data requirements, executable trajectories, and open-ended answers grounded in execution evidence. With a shared GPT-5 backbone, Earth-Agent-Pro achieves 66.13% LLM-as-Judge accuracy, exceeding ReAct by 20.95 points in this metric and 24.44 points in Tools-In-Order. Joint adapter tuning raises Qwen3.5-9B LLM-as-Judge accuracy from 38.31% to 50.00%, an 11.69-point gain over the untuned configuration. Planning-only evaluation and execution with the reference workflow show that the adapters improve workflow composition and argument grounding, respectively. Code and datasets will be released soon.

Fri 11 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language
The gist
Earth observation involves collecting data from satellites and sensors to answer big science questions about our planet. The authors developed Earth-Agent-Pro, a system that can take those big questions and turn them into step-by-step plans to gather and analyze data automatically. Their system uses techniques to remember and fix parts of plans if new information changes things. They tested it on many tasks involving satellite images and other data, and it performed better than previous methods.
Open 2609.12533v1

AI predicts flare stack combustion efficiency using thermal video

AI-Powered Flare Combustion Efficiency Estimation

Abstract: Achieving high combustion efficiency in flare stacks is crucial for adhering to regulatory standards and controlling the release of hydrocarbons into the environment. Traditional instruments like gas analyzers and hyperspectral cameras are expensive, fragile, and require frequent calibration, which makes them impractical for remote or budget constrained industrial sites. We propose an innovative solution that combines a lightweight vision-language encoder with a compact multi-layer perceptron to predict combustion efficiency directly from low-cost thermal video footage. The fully trained model is integrated into an easy-to-deploy graphical user interface. This interface overlays predicted combustion efficiency values on each video frame, displays real-time trends in combustion efficiency, shows the distribution of combustion efficiency across all frames in the video, and allows users to export CSV reports. Over a six-month period, the system achieved 99% uptime and required less than 15 minutes of maintenance per week.

Thu 10 SeptArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Burning hydrocarbons in flare stacks needs to be efficient to meet regulations and reduce pollution. The authors created an AI system that uses thermal videos, which are cheaper and easier to get than traditional instruments, to estimate how well the flare burns. Their system displays this information in real time with easy-to-understand visuals and needed very little maintenance over six months. This approach helps industries monitor flares safely and more affordably.
Open 2609.11262v1

Vision language models tested on coastal and underwater scenes

Evaluation of Vision-Language Models Across Diverse Coastal Environments

Abstract: Vision-language models (VLMs) enable robotic per- ception by associating visual observations with natural-language concepts. Yet their performance in coastal environments remains largely unexplored. We introduce a densely labeled coastal dataset containing more than 1,000 images collected across seven missions in three regions of Oahu, Hawaii, with 18 semantic classes and over 7,400 annotated instances. We evaluate seven modern VLMs through three complementary experiments mea- suring text-to-mask, mask-to-mask, and mask-to-text alignment. Broad landscape classes are generally recognized more accurately than conventional object and coastal classes, with coastal con- cepts presenting the greatest challenge. However, comparisons of shared conventional classes across coastal and terrestrial datasets reveal no consistent performance difference attributable solely to environmental context. Mask-to-mask matching also remains similar across conventional and coastal classes, while alternative textual labels substantially improve recognition of several coastal concepts. These results suggest that lower performance on coastal classes (at least on the objects/query categories evaluated) is heavily influenced by segmentation and linguistic representation.

Wed 9 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Vision-language models help robots understand pictures by linking images to words. This paper studies how well these models work with photos from coastal areas in Hawaii, which can be different from land scenes. The authors made a special dataset of over 1,000 coastal images with detailed labels and tested seven popular models. They found that models recognize broad landscapes better than specific coastal objects, and special labels can help improve recognizing some coastal things. The challenges come mainly from how objects are segmented in images and described by words.
Open 2609.10855v1

Prompt guidance improves weak water mapping in multispectral images

Beyond Weak Labels: Prompt-Guided Local Refinement for Weakly Supervised Water Segmentation in High-Resolution Multispectral Imagery

Abstract: High-resolution water mapping supports environmental monitoring and related applications, but accurate pixel-level labels are difficult and costly to produce. Official hydrographic vectors provide scalable weak supervision, but they contain artifacts like boundary noise, temporal mismatch, and omissions of small water structures. We propose a two-stage framework for weakly supervised water segmentation in high resolution multispectral imagery. Stage 1 learns initial masks from rasterized vector pseudo-labels, and Stage 2 converts these masks into structured component-wise prompts for localized refinement. On a manually corrected validation set, refinement improves SegFormer-B0 from 0.9509 to 0.9535 IoU and U-Net from 0.9408 to 0.9486 IoU, with corresponding F1 gains from 0.9749 to 0.9762 and 0.9695 to 0.9736. It leads to sharper shorelines, reduced boundary spillover, and better thin-structure delineation. The results indicate that prompt-guided refinement can improve pseudo-label-based water segmentation by targeting local errors that are poorly captured by global training supervision.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Mapping water areas accurately in high-resolution satellite images is important but hard because detailed water labels are costly to make. The authors use rough water boundary data as a starting point, then improve these guesses by focusing on local areas that are often wrong. Their two-step method sharpens shorelines and catches small water details better. This makes water maps from satellite pictures more precise without needing expensive manual labels.
Open 2609.10371v1

Pca and random forest improve hyperspectral image classification accuracy

Dimensionality Reduction for Hyperspectral Image Classification

Abstract: This paper addresses the issue of supervised classification in the context of hyperspectral satellite images. It deals with two fundamental aspects: dimensionality reduction of data and the selection of appropriate supervised classification techniques. Firstly, we delve into dimensionality reduction, a critical step in simplifying the management of hyperspectral data. The reduction aims to decrease complexity in terms of memory and computing time. We examine two commonly used methods: Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA). Subsequently, we explore the selection of the most suitable supervised classification algorithms for hyperspectral images. We compare the performance of three methods: K-Nearest Neighbors (KNN), Support Vector Machines (SVM), and Random Forest (RF) using real hyperspectral data. The results highlight that the combination of PCA and RF yields the highest overall accuracy and Kappa coefficient.

Wed 9 SeptComputer Vision and Pattern Recognition
The gist
Hyperspectral images contain lots of data that can be hard to manage and analyze. The authors look at ways to reduce this data without losing important details, using techniques like PCA and LDA. They also compare different computer methods to classify the images. Their tests show that using PCA together with a Random Forest classifier gives the best results for identifying features in satellite images.
Open 2609.10334v1

Methane detection improves by fusing multiple satellite data sources

MethaneFuse: Learning from Multi-Sensor Satellite Observations for Methane Plume Detection

Abstract: Methane plume detection from satellite imagery is constrained by incomplete observations: public satellites provide complementary spatial, spectral, and atmospheric evidence, but real plume cases rarely contain fully paired multi-sensor measurements because of revisit schedules, cloud coverage, acquisition quality, and the transient nature of emissions. Most learning-based detectors rely on single-sensor inputs, especially Sentinel-2 (S2), leaving many reported plume cases unusable. We construct MethaneUnion, a temporal multi-sensor dataset built from Carbon Mapper plume reports and matched S2, Landsat 8/9 (L8/9), EMIT, and Sentinel-5P (S5P) observations. Built on MethaneUnion, MethaneFuse learns from heterogeneous satellite observations under partial sensor availability without requiring complete four-sensor measurements. MethaneUnion expands usable coverage from 3,211 valid S2-matched plume cases to 8,981 reported plume cases with multi-sensor observations. At the representative 480 m setting, MethaneFuse achieves 84.87 F1 and 93.62 AUROC, improving over the strongest baseline by 5.65 F1 and 8.30 AUROC points while reducing false positives by 8.19 points. Sensor-availability experiments show that MethaneFuse improves detection when S2 is available and transfers plume knowledge to L8/9, EMIT, and S5P when S2 is unavailable. These results demonstrate the value of learning from incomplete heterogeneous sensor observations for practical methane plume detection.

Wed 9 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Detecting methane gas leaks from space is tricky because satellites don’t always capture complete data at once due to factors like clouds and timing. The authors built a new dataset combining reports of methane plumes with images from several satellites, even when not all data types are present at the same time. They created a method called MethaneFuse that learns to spot methane plumes using this mixed, incomplete data. This approach improves detection accuracy and reduces false alarms compared to previous methods using only single satellite data.
Open 2609.09762v1

Hyperbolic geometry improves unknown object detection in satellite images

Hyperbolic Geometry for Open-World Object Detection in Remote Sensing Imagery

Abstract: Open-world object detection (OWOD) extends closed-set detection by requiring models to identify unknown objects and incrementally learn them once annotations become available. In remote sensing imagery, object categories often exhibit latent hierarchical relationships that may be inadequately represented in the Euclidean spaces commonly adopted by existing methods, limiting unknown-object recall and incremental-learning performance. To address this issue, we investigate hyperbolic geometry for OWOD in remote sensing imagery and propose HyRS-OWOD. To improve unknown object recall, we design a two-step unknown-object discovery mechanism: a Decoupled Objectness Learning (DOL) module that disentangles foreground perception from semantic information to separate foreground proposals from background regions, followed by a Hyperbolic Uncertainty Learning (HUL) component that leverages the radius of hyperbolic embeddings as an uncertainty-aware cue for known-unknown discrimination. For incremental learning, we develop a Hyperbolic Metric Learning (HML) strategy that enhances inter-class separability, facilitating the incorporation of novel categories while mitigating catastrophic forgetting. Experiments on three remote sensing benchmarks demonstrate consistent improvements in unknown recall and incremental learning over state-of-the-art OWOD methods.

Wed 9 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Detecting new or unknown objects in satellite images is a challenge because objects often have hidden relationships that usual methods miss. The authors use hyperbolic geometry, a type of math that better represents these relationships, to help identify unknown objects more accurately. They built a system that separates objects from backgrounds and measures uncertainty to decide if an object is known or new. Their method also helps computers learn new object types over time without forgetting old ones. Tests on satellite image datasets showed their approach works better than existing methods.
Open 2609.09626v1

Vision language models better identify agriculture problems with rubric guidance

Vision-language models know more about agriculture than they show and rubric-grounded verifications close the gap

Abstract: Vision-language models (VLMs) show promise for agricultural classification, but zero-shot performance on disease, pest, damage, quality, and species identification remains poor, and it is unclear whether this reflects weak visual features or a failure to connect them to domain knowledge. We build a benchmark of 116 datasets, 834 classes, and 8,324 images spanning these tasks to isolate where the gap arises. Linear probing shows VLM vision encoders already encode agricultural features nearly as separable as a self-supervised DINOv3 baseline, ruling out weak visual representations as the primary bottleneck. Conditioning each model on an oracle reference description (an upper bound on its parametric knowledge) closes most of the gap left by an unaided lower bound, showing VLMs already know more about agriculture than they show. To close this gap without an oracle description at inference time, we structure test-time reasoning around a fixed, per-task diagnostic rubric: the model generates $K$ candidate responses and a Probabilistic Pivot Tournament (PPT) verifier, scored pairwise against the rubric, selects the best one. This nearly doubles judged F1 over the lower bound and matches or exceeds the upper bound on several tasks, notably pushing Gemma 4 E4B-it's disease F1 to 0.71, above its own upper bound of 0.60. However, the verifier's letter-scale confidence score has the opposite of its intended effect: filtering to its most confident predictions does not improve accuracy and correlates negatively with correctness across every model and pool size tested, so the score cannot serve as a measure of predictive uncertainty, and most of the observed gain likely comes from rubric-grounded generation rather than pairwise verification.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Vision-language models struggle to identify agricultural issues like diseases and pests without extra help. The authors found that these models actually recognize important visual features but don’t always connect them to the right agricultural knowledge. By giving the models a clear set of guidelines to check their guesses, the team improved their accuracy significantly, showing these models know more than they initially appear to. However, the confidence score the system uses doesn’t reliably tell when the model is correct.
Open 2609.09417v1

Arctic navigation tool balances safety ecology and community needs

An Autonomous GeoAI Agent for Arctic Eco-Navigation

Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria problem: routes that improve vessel safety or efficiency may increase exposure to sea ice, sensitive ecosystems, or nearby communities. Existing routing methods prioritize travel time, fuel use, and navigational risk, often overlooking ecological and community impacts. We introduce a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, physical, ecological, and community-related criteria within a unified routing framework. Multiple specialized agents coordinate geospatial data acquisition and preparation, multi-objective route generation, and skyline-based decision support. The ecological criteria explicitly account for exposure to sensitive areas, including Essential Fish Habitat and seal critical habitat. By considering these ecosystem impacts and potential community burdens while keeping consequential value judgments under human control, the framework supports safer, more transparent, and socially responsible Arctic navigation. Project page and code are publicly available. https://samiraat.github.io/Arctic-Eco-Navigation-Agent/, https://github.com/samiraat/Arctic-Eco-Navigation-Agent

Tue 8 SeptArtificial Intelligence
The gist
Navigating ships through the Arctic is tricky because of ice, environmental concerns, and local communities nearby. The authors made a computer system that helps plan safer routes by looking at many factors like ice conditions, wildlife habitats, and community impacts all at once. This system uses different specialized parts working together and lets people make final choices, making sure decisions consider the environment and communities responsibly. The tool is openly shared online so others can use or improve it.
Open 2609.09374v1

New method improves accuracy of hyperspectral image material detection

Interpretable Hyperspectral Unmixing Framework with Fixed Endmember Prior and Structured Residual Refinement

Abstract: Hyperspectral unmixing decomposes mixed pixels into material endmembers and their abundances from contiguous spectral observations. In modular sensing pipelines, endmembers are often first identified and then treated as fixed during abundance estimation. When this fixed endmember prior is inaccurate, spatially structured mismatch arising from illumination changes, sensor artifacts, or material boundaries may be incorrectly captured by the abundance variables, leading to unstable decompositions. This study presents an interpretable stage-wise hyperspectral unmixing framework (I-HyperSU) under fixed endmember priors, which is explicitly decomposed into a fixed endmember matrix $\mathbf{A}$, an abundance block $\mathbf{X}$, and a structural residual refinement block $\mathbf{S}$. The X-block estimates abundances using FISTA with nonnegativity and sparsity enhancement, and a soft penalty that approximately enforces sum-to-one constraints. The S-block jointly applies low-rank SVD structural regularization and a lightweight deep image prior (DIP) to refine structured residuals. This staged design makes the interaction between abundance and residual components transparent and interpretable. Experiments on Samson, Urban, and Jasper Ridge datasets demonstrate that, under fixed and imperfect endmember priors, soft abundance relaxation consistently outperforms hard simplex projection. Under the default N-FINDR endmember prior, the proposed framework reduces the joint reconstruction error by 61.7\%--69.5\% compared with a fixed-$\mathbf{A}$ UCLS baseline, while keeping the abundance RMSE nearly unchanged, indicating that the residual refinement branch accounts for structured model mismatch without degrading the abundance estimates. For example, on Urban, the reconstruction SAM decreases from $5.99^\circ$ for the X-only model to $1.92^\circ$ for the full model.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Hyperspectral images show many colors for each pixel, helping identify materials, but these pixels often mix several substances, making it hard to tell them apart. The authors introduce a step-by-step method that treats the known materials as fixed and better separates the mixture by adding a correction step for errors caused by lighting or sensor problems. Their approach improves the accuracy of separating materials without messing up the estimated amounts, making the results easier to understand. Tests on standard datasets showed much better image reconstruction and clearer recognition of materials.
Open 2609.08786v1

Transformer models improve activated sludge image recognition accuracy

Effects of model architecture and learning strategies on deep learning-based recognition of activated sludge microscopic images and comparison with quantitative image analysis

Abstract: Microscopic image analysis has long been recognized as a promising approach for monitoring activated sludge. In recent years, deep learning-based image analysis has been increasingly adopted in this field because of its high performance. However, previous studies on microscopic image analysis of activated sludge have rarely explored transformer-based models or self-supervised foundation models and have instead relied on CNNs and supervised ImageNet pretraining. In addition, previous studies often downsampled image sizes, but the effects of downsampling have not been sufficiently investigated, and the relationship between downsampling strategies and image analysis performance remains unclear. Furthermore, no study has quantitatively compared deep learning performance with quantitative image analysis (QIA), which was widely used before the emergence of deep learning. In this study, to examine how model architecture and learning strategies affect performance in microscopic image analysis of activated sludge and to quantitatively determine whether deep learning outperforms QIA, we prepared three types of activated sludge samples, classified their microscopic images, and evaluated classification accuracy. Our results showed that transformer-based architectures and alternative pretraining methods were effective in terms of classification accuracy. Our downsampling analysis showed that using overly small images reduced accuracy, but increasing image size beyond a certain point did not improve it further. In addition, the analysis indicated that, to achieve high classification accuracy, maintaining the field of view was a more effective downsampling strategy than maintaining resolution. Finally, our comparison between deep learning and QIA showed that deep learning outperformed QIA in terms of accuracy.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Identifying what’s in sludge under a microscope helps keep water clean, but it’s tricky. The authors tested different computer methods to recognize sludge images and found that newer transformer-based models worked better than older ones. They also learned that keeping the overall scene in images matters more for accuracy than just zooming in on details. Finally, their advanced models beat older image analysis techniques in correctly identifying sludge types.
Open 2609.08570v1

Model chaotic physical processes directly from time series data

Equation free data-driven modelling of chaotic processes

Abstract: We introduce a method for constructing predictive models of non cyclic physical processes directly from time-series data, without assuming an underlying differential equation. The observations define a discrete evolution rule whose recurrent behaviour captures the essential dynamics of the process. Analysing this behaviour across multiple geometric scales leads to probabilistic models in the form of Markov chains. Hyperbolicity criteria identify when these models provide a consistent statistical description of the data. The method is inspired by, and illustrated through, the analysis of a biological imaging data set referred to as the Cell Process.

Tue 8 SeptInformation Theory
The gist
Predicting complex and irregular natural events, like certain biological processes, is difficult because they don't follow simple repeating patterns. This paper presents a way to build models directly from sequences of measurements, without needing to know the exact equations behind the process. The method uses patterns in the data's behavior across different scales to create statistical models that can describe how the system evolves. The authors show how their approach works using biological imaging data.
Open 2609.08530v1

Remote sensing models select image regions to track changes over time

From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMs

Abstract: Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging. Existing approaches rely on generating bounding box coordinates as token sequences, which is fragile for the small, densely packed objects common in remote sensing and increasingly error-prone when multiple targets must be localized simultaneously. In this work, we present an RS-specific formulation of the region selection paradigm, previously explored in natural-image MLLMs, and extend it to temporal change localization over multi-image sequences. Our framework employs a text-conditioned region proposal module, encodes each candidate as special tokens carrying per-frame visual features enriched with spatial and temporal cues, and lets the LLM localize targets by selecting region tokens in its response. We construct a multi-task training and evaluation suite spanning localization, referring expression, visual grounding, and understanding tasks across single-image and multi-temporal settings. Experiments show that our approach substantially outperforms coordinate-generation baselines on temporal change localization, while improving single-image visual grounding and maintaining competitive understanding performance. Oracle analysis decomposes the contributions of the region proposer and the LLM selector, providing diagnostic insight unique to this framework. Our code will be available at https://github.com/juwan-kr/RS-RegionSelect.

Tue 8 SeptComputer Vision and Pattern RecognitionComputation and Language
The gist
Identifying exactly where things have changed in satellite images is hard, especially when many small objects are close together. The authors developed a new way for computers to pick out specific areas instead of just giving coordinates, helping the model find changes across multiple images taken over time. Their tests show this method is more reliable for spotting changes while still understanding single pictures well. They also analyze how different parts of the system contribute to its success, making it easier to improve in the future.
Open 2609.08391v1

New low-rank model improves hyperspectral anomaly detection speed and flexibility

Hyperspectral Anomaly Detection via Group Sparse Low-Rank Tensor Factorization With Automatic Anomaly Grouping

Abstract: Low-rank tensor modeling has become an effective tool for hyperspectral anomaly detection. However, existing methods still suffer from high computational cost and limited flexibility in characterizing spatially structured anomalies. To address these issues, this paper proposes a hyperspectral anomaly detection method based on group sparse low-rank tensor factorization with automatic anomaly grouping (GSAA). Specifically, the low tubal rank background is characterized by imposing group sparsity on tensor factors, which provides an efficient alternative to direct tensor rank regularization. For anomaly modeling, a latent grouping map is introduced to build an automatic anomaly grouping penalty, allowing anomaly groups to be adaptively inferred from the data rather than predefined at the pixel level. To further exploit complementary spectral and spatial information, GSAA is applied in both domains, and the resulting detection maps are fused to form a spectral--spatial version of GSAA, termed GSAA-SS. An efficient linearized alternating direction method of multipliers algorithm with convergence guarantee is developed to solve the resulting model. Experimental results on five real hyperspectral datasets demonstrate that the proposed method achieves superior detection performance and competitive computational efficiency compared with several state-of-the-art methods.

Tue 8 SeptComputer Vision and Pattern Recognition
The gist
Detecting unusual materials or objects in hyperspectral images is hard because existing methods are slow and don’t handle groups of anomalies well. The authors propose a method that simplifies the math using group sparsity, making detection faster and better at finding clusters of anomalies automatically. They combine information from both the spectral details and the image layout to improve results. Their approach works better than current methods on several real-world hyperspectral datasets.
Open 2609.08121v1

Risk-aware experimental design improves prediction for rare outcomes

Risk-Aware Goal-Oriented Bayesian Optimal Experimental Design

Abstract: Traditional Bayesian optimal experimental design (OED) selects measurements that best inform a model's parameters. However, such measurements can be suboptimal for downstream predictions. Goal-oriented OED targets the prediction directly. However, the existing goal-oriented criteria value all reductions in predictive uncertainty equally, with no way to prioritize rare, high-consequence outcomes. In this article, we develop a risk-aware framework that composes risk at three levels, each generalizing an ingredient of classical $I$- and $G$-optimal design: a deviation measure of the posterior predictive uncertainty (generalizing the predictive variance), a risk measure across the prediction domain (interpolating $I$-optimal averaging and $G$-optimal worst-case selection), and a risk measure over datasets (generalizing the expectation). We generate each level from a regret function in the risk quadrangle, so that one triple specifies a practitioner's risk preference. We relax the design to continuous weights on the unit simplex and construct a nested-quadrature estimator that is differentiable in the design variable. This enables solving the optimal design problem with gradient-based methods, avoiding a combinatorial search over candidate designs. For a linear-Gaussian lognormal model and a nonlinear extension, we derive closed-form objectives. These give exact references against which we verify that the estimator converges. We demonstrate this framework for finding optimal sensor placements in an inverse problem governed by an advection-diffusion equation. We find that the risk-aware designs substantially outperform the expected-information-gain baseline, which is statistically indistinguishable from a random allocation.

Mon 7 SeptComputational Engineering, Finance, and Science
The gist
Traditional methods pick measurements to better understand model details but don’t always help with important predictions. The authors propose a new approach that focuses on measurements that improve predictions, especially for rare but high-risk events. Their method lets users specify what kinds of risks matter most and uses math techniques to efficiently find the best measurement plans. They tested this on a problem involving sensor placement for tracking how substances spread and showed it works better than standard methods.
Open 2609.08023v1

Semi supervised learning struggles with biased spatial data sampling

Semi-Supervised Learning under Spatially Biased Sampling

Abstract: Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as three distinct mechanisms, varied independently via a labelled-sampling concentration parameter, a spatial length scale, and a non-stationarity strength parameter, and ask how mismatch degrades SSL, whether the cluster and manifold assumptions survive it, and how the resulting failure can be diagnosed. Using a controlled synthetic framework alongside PovertyMap-WILDS, California housing, socio-economic and US air quality monitoring datasets, we systematically vary the degree of mismatch while accounting for spatial autocorrelation and non-stationarity. Through a series of analyses including a segmented-regression changepoint, we show that in the synthetic generator, SSL performance does not degrade gradually but instead exhibits a threshold-like breakdown between approximately 0.71 and 0.77 once distribution mismatch becomes sufficiently severe. We further demonstrate that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available. To support practical deployment, we evaluate several distribution-divergence measures as indicators of reliability and introduce a kernel-weighted local divergence metric that provides a more stable estimate of spatial mismatch than a naïve localised approach. These findings provide empirical evidence and diagnostic tools for better documenting the risk of incorporating unlabelled spatial data into semi-supervised learning workflows.

Mon 7 SeptMachine Learning
The gist
Semi-supervised learning usually assumes that labelled and unlabelled data come from the same overall pattern. The authors found that when labels come from biased locations in space, this assumption breaks down sharply rather than gradually. They also show that spatial changes in data traits and correlations affect learning independently. Their work introduces ways to detect when this mismatch happens, helping users know when unlabelled spatial data might harm learning performance.
Open 2609.07982v1