Graph neural network-based multimodal sensor fusion for robust autonomous driving perception
Yuan Wei · 2025
Multimodal sensor fusion is a fundamental component of autonomous driving perception, combining data from LiDAR, radar, and cameras to provide a comprehensive understanding of dynamic environments. Nonetheless, traditional fusion methods face challenges such as sensor noise, temporal misalignment, and environmental complexities, which often limit their robustness and scalability. This study introduces a novel Graph Neural Network (GNN)-based framework that leverages graph representations to explicitly model spatial, temporal, and semantic relationships across modalities. Sensor features are encoded as graph nodes, while edges capture proximity and motion coherence, enabling effective feature aggregation through a message-passing mechanism. Extensive experiments on public datasets, including KITTI and nuScenes, demonstrate that the proposed framework achieves significant improvements in object detection and semantic segmentation tasks. Specifically, the GNN-based approach outperforms traditional early, mid, and late fusion methods, achieving higher accuracy, F1-scores, and resilience under adverse conditions such as rain, fog, and occlusions. The proposed framework successfully integrates the complementary strengths of multimodal sensors, offering a robust and scalable solution for autonomous driving perception. These findings underscore its potential to enhance the accuracy and reliability of perception systems, paving the way for advancements in intelligent autonomous driving technologies.