CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
2026-08-10 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors address the problem of combining data from different sensors in self-driving cars, like cameras and radar, which can be unreliable in bad weather or low visibility. They propose a method called CRUISE that uses a vision-language model to better estimate how uncertain each sensor input is at a detailed, pixel-by-pixel level. This allows the system to decide which sensor readings to trust more when fusing the data together. They also add a mechanism to capture how different sensors relate to each other, improving overall reliability. Their approach aims to make sensor fusion smarter and more adaptable in challenging conditions.
autonomous vehiclessensor fusionuncertainty quantificationvision-language modelLiDARcross-modal fusionpixel-level uncertaintyadaptive mechanismout-of-distribution generalization
Authors
Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque
Abstract
Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to prioritize reliable signals, existing uncertainty-aware fusion methods typically rely on simple feature-level uncertainty estimates and thus often fail to generalize effectively in complex, out-of-distribution scenarios. To address this limitation, we propose CRUISE, a novel uncertainty-aware cross-modal sensor fusion framework. CRUISE integrates a vision-language model (VLM)-guided UQ module that generates fine-grained, pixel-level uncertainty estimates. By leveraging the VLM's rich prior knowledge and superior contextual reasoning, our approach provides a highly informative guide for the fusion process. Furthermore, we introduce a dynamic adaptive mechanism that explicitly models and captures cross-modal dependencies, ensuring the framework fully exploits the inherent complementary nature of multi-sensor inputs.