Multi sensor fusion improves vehicle network beam prediction accuracy
Robust Beam Prediction for V2X Networks with Multi-Modal Sensing
Machine Learning
Summary
In vehicle communication networks, it's important to focus signals accurately to maintain good connections. The authors noticed current methods rely mostly on radio signals, which can fail in tricky environments. They designed a system called BeamTransFuser that combines data from cameras, LiDAR, radar, and GPS to better guess where to direct signals. Their system also copes when some sensors aren’t available by filling in missing data from the others. Tests on real vehicle data showed this approach predicts signal directions more reliably than before.
What this means in practice
- •For automotive communication engineers: Design vehicle networks that maintain strong connections even when some sensors fail by fusing multiple sensor data for beam prediction.
- •For smart city infrastructure teams: Improve roadside unit communication by using multi-modal sensing to enhance beamforming accuracy in complex urban settings.
Authors
Chen Shang, Dinh Thai Hoang, Diep N. Nguyen, Jiadong Yu
Abstract
Integrated sensing and communication (ISAC) provides a promising foundation for beam prediction in future vehicle-to-everything (V2X) networks. However, existing sensing-assisted beamforming methods still rely heavily on radio-frequency sensing, which may become unreliable in complex vehicular environments. Meanwhile, the growing availability of heterogeneous sensors, such as cameras and LiDAR, offers new opportunities to improve beam prediction through richer environmental perception. Motivated by this, this paper proposes a multi-modal beam prediction framework for V2X networks. Specifically, we develop BeamTransFuser, a hierarchical Transformer-based architecture that progressively fuses camera, LiDAR, radar, and GPS observations for robust beam prediction. In addition, to handle possible missing modalities in practical deployment, we introduce a generative module that reconstructs missing modality features from the available observations. Experimental results on a real-world multi-modal V2X dataset show that the proposed framework consistently outperforms representative baselines, while the generative module further improves robustness under incomplete sensing conditions.