Foundation model confidence and depth improve detection of unknown objects

Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation

Computer Vision and Pattern Recognition

Summary

Self-driving cars need to notice when they see things they haven't seen before, like strange animals or loose items on the road, to stay safe. The authors propose a way to detect these unusual objects without needing extra training or special data. They combine how confident a big AI model is about what it sees with depth information from a single camera to spot these unknown things more reliably. Their method works well on road-focused tests and in video sequences, helping cars better understand their surroundings.

What this means in practice

Authors

Serin Varghese, Fabian Hüger, Kira Maag

Abstract

Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.