Dynamic Scene Understanding for Autonomous Driving Using 2D-3D Convolution With Voxel Key Points
Kunhua Liu, Yi Zheng, Junkun Xie, Yuting Xie, Feiyang Wang, Longyu Ma, Chenggang Dai, Tao Lu · IEEE Transactions on Intelligent Transportation Systems · 2024
With the growing emphasis on real-time 3D data processing in autonomous driving, robotics, and intelligent vehicles, the demand for efficient point cloud processing has expanded significantly. Early deep learning approaches to point cloud semantic segmentation relied on volumetric grids and projections, which often compromised the inherent geometric structure of point clouds. More recent methods attempt to learn directly from raw point clouds, focusing on local neighborhood information; however, optimizing computational efficiency for dynamic scenes remains a challenge. This paper presents a novel 2D-3D convolutional framework, VKPNet, for point cloud semantic segmentation that leverages Voxel Key Points (VKPs) to efficiently aggregate local features and enhance receptive fields. The proposed approach first initializes 3D point cloud features using 2D image features and applies a heuristic method to filter 3D points, extracting only those necessary for semantic segmentation, thereby reducing the input data scale. VKPs are introduced to aggregate local features at voxel cube vertices, and a 3D convolution based on VKPs is designed to expand the receptive field, facilitating effective spatiotemporal feature learning. Experimental results on the ScanNet and Semantic KITTI datasets validate the effectiveness of our VKPNet model. The framework achieves mIoU scores of 0.735 on the ScanNet dataset and 0.689 on the Semantic KITTI dataset, with a processing speed of 0.09 seconds per frame on ScanNet. These results demonstrate that VKPNet not only outperforms prior methods across various benchmarks but also achieves efficient and accurate semantic segmentation in dynamic scenes.