Collaborative perception improves model adaptation in self driving
Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation
Computer Vision and Pattern RecognitionMachine Learning
Summary
When self-driving cars move to new places, their computer vision systems often make mistakes because they haven't seen those environments before. The authors present a way to improve these systems by letting multiple cars share information and learn from each other's views, which usually gives better clues than just one car alone. They also design smart methods to share only the most important data, avoid confusing information when views don't match, and gradually use these shared clues to update the model. Their tests show this approach works better than previous techniques that only use data from one car.
What this means in practice
- •For autonomous vehicle system integrators: Improve perception accuracy in autonomous cars by using shared sensor data from multiple vehicles to update models without manual labeling.
- •For surveillance system operators: Enhance object detection across distributed camera networks by adapting models using collaborative views for more reliable monitoring in new environments.
Authors
Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, Yuguang Fang
Abstract
In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts. While unsupervised model adaptation offers a feasible solution without labor-intensive manual labeling, existing methods that rely solely on the ego-vehicle's data often lead to inferior pseudo-labeling performance. To address this critical issue, we propose LDE, Learning from Distributed ``Eyes", a novel framework that transforms collaborative perception (CP) into a source of high-quality supervision for model adaptation. This pseudo-labeling approach is hyperparameter-insensitive and relatively reliable, assuming CP often outperforms single-agent's perception. However, naively implementing this approach encounters (1) the communication bottleneck of sharing rich features under time and bandwidth constraints, (2) the view discrepancy between the CP view and the learner's Field of View (FoV), and (3) the unreliability even in CP-generated labels. To address these issues, we design an adaptation-oriented feature sharing mechanism that selectively transmits the most critical information for adaptation, an FoV filtering method that meticulously eliminates mismatched labels, and a curriculum learning strategy to progressively exploit pseudo labels. Extensive experiments on 3D object detection tasks demonstrate that LDE consistently outperforms both the pre-trained models and state-of-the-art unsupervised adaptation methods.