End-to-End Multi-View Fusion for Enhanced Perception and Motion Prediction
Yasser H. Khalil, Hussein Talaat Mouftah · 2021 IEEE 94th Vehicular Technology Conference (VTC2021-Fall) · 2021
Perception and motion prediction are indispensable components to the smooth operation of autonomous vehicles and the safety of the surrounding environment. Strengthening the accuracy of perception and motion prediction in autonomous vehicles remains of paramount importance. Therefore, we propose an end-to-end multi-view fusion methodology applied to MotionNet backbone network to enhance the sharpness of both perception and motion prediction. MotionNet is a state-of-the-art real-time model designed for joint perception and motion prediction. Our multi-view input is based on a single LIDAR sensor and formed by the fusion of range view features with bird's eye view. We evaluate our proposed work on nuScenes dataset and demonstrate through experiments that our proposed extension to MotionNet using the multi-view fusion technique outperforms MotionNet in both perception and motion prediction, especially for small and distant objects.