Preference guidance improves open-vocabulary image segmentation in new domains
Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
Computer Vision and Pattern Recognition
Summary
Semantic segmentation assigns labels to every pixel in an image, but models often struggle when applied to specialized areas like medical images or satellite photos. This paper shows how asking simple yes-or-no questions about uncertain parts of an image can help improve these models without needing detailed, hard-to-get labels. The authors use disagreements between different text prompts as clues to figure out where to ask these questions, leading to better segmentation results. Their method works well on various existing models and remains effective even if some answers are noisy.
What this means in practice
- •For medical image analysts: Improve image segmentation models in medical scans using simple binary feedback instead of costly pixel-level annotations.
- •For remote sensing teams: Enhance segmentation accuracy of satellite or aerial images by querying local uncertainty with minimal expert input.
Authors
Hyun-Kurl Jang, Jihun Kim, Kuk-Jin Yoon
Abstract
Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences. Our code is available at https://github.com/blue-531/pref-ovss.