Integrating sensing and communication (ISAC) is a promising technology for predictive beamforming in 6G vehicle-to-everything (V2X) networks. However, current ISAC paradigms rely solely on radio-frequency (RF)-based sensing, which limits sensing resolution and beamforming robustness in complex wireless environments. Fortunately, the widespread deployment of diverse non-RF sensors such as cameras and LiDAR, along with the integration of artificial intelligence (AI) and communication systems, offers new opportunities to improve the synergy between sensing and communication. Motivated by this, this work develops a multi-modal sensing-assisted beamforming framework for realistic V2X scenarios. Specifically, we propose BeamTransFuser, a hierarchical Transformer-based multi-modal learning framework that exploits cross-modal correlations among camera, LiDAR, radar, and GPS observations to improve beam prediction accuracy and robustness. To facilitate practical deployment on roadside units, we further develop a module-aware pruning scheme to reduce inference latency while preserving competitive performance. Furthermore, to address potential missing-modality conditions in real-world scenarios, we introduce a generative model that is able to reconstruct missing inputs from available observations, allowing the framework to operate reliably even under incomplete sensing conditions. Extensive experimental results conducted on real-world datasets demonstrate that the proposed scheme consistently outperforms existing baselines across various metrics.
Paper
References (50)
Scroll for more · 38 remaining