Aperture improves remote sensing image interpretability without training
Aperture: Training-Free Multiscale Concept Bottlenecks for Remote Sensing
Computer Vision and Pattern Recognition
Summary
Models that look at satellite images often struggle to explain their decisions clearly. The authors created Aperture, a model that does not need training but can identify detailed features in images using a special method to focus on small areas and concepts at different scales. They tested it on a new detailed dataset across three countries and found it performed better than other similar models, even those that require training. This helps experts understand what the model sees and why it makes certain classifications.
What this means in practice
- •For geospatial analysts: Enhance interpretability of remote sensing classifications without costly model retraining by using Aperture’s multiscale concept bottleneck approach.
- •For urban planning teams: Use detailed concept maps derived by Aperture to monitor small-scale urban changes across regions with improved reliability and minimal annotation effort.
- •For environmental monitoring companies: Offer training-free interpretability tools for satellite image analysis products that improve client trust by showing clear concept-level reasoning.$Commercial implications: Enables development of interpretable satellite image analysis software that clarifies decisions, appealing to clients needing explainable environmental insights.
Authors
Rishabh Mondal, Nipun Batra, Utkarsh Mall
Abstract
While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability and expert interaction, they are either too expensive to train for the remote sensing domain or perform poorly without annotation. We posit that in expert domains like remote sensing, such training-free models require both fine details in both image and concept space. In image space, we propose a multiscale concept bottleneck using greedy quadtree routing to locate small concepts. In concept space, we replace contrastive vision language models with pre-trained MLLMs and present a way to get reliable concept scores from them. We introduce APERTURE that blends concept scores at the global image and native concept-scale level to give state-ofthe-art training-free model performance. To test these models, introduce SiFC, a fine-grained concept-centric dataset across three countries, with human-reviewed class-level concept maps. On SiFC, APERTURE outperforms the best training-free baselines by more than 10 percentage points in macro F1-score, and notably also outperforms supervised concept bottleneck models. Targeted component-removal tests examine whether concept scores respond to changes in visual evidence, while temporal experiments show that descriptor updates improve recognition of technological changes without retraining.