PVF-NET: Point & Voxel Fusion 3D Object Detection Framework for Point Cloud

Zhihao Cui, Zhenhua Zhang · 2020

In this paper, we present a novel 3D object detection framework for locating 3D bounding boxes of the target in autonomous driving scenes. Our proposed framework consists of two novel modules, which are twofold proposal fusion module and the RoI deep fusion module. In the former module, we utilized the 3D voxel Sparse Convolution Neural Network (CNN) and PointNet-like network to coarse generate the voxel-based and point-based proposals, where these proposals contain voxel-dense and point-wise features under the raw point cloud. Twofold proposal fusion module integrated those proposals and extended the proposal generalization, thereby dramatically improve the proposals’ recall rate in the first stage for further utilizing in the proposals refinement stage. Given the coarse integrated 3D proposals produced by the twofold proposal module, the RoI deep fusion module is proposed to abstract and aggregate the multi-scale voxel-based feature and the point-wise feature through voxel-aware pooling layer and point-aware pooling layer, respectively. Follow by that, the specific features on the different proposals are integrated via the proposals-aware fusion layer to further enrich the feature dimensionality and utilize the high-quality proposals for the proposals refinement stage to reinforce the prediction of the target bounding boxes. We conduct the experiments on KITTI dataset and evaluate our method on 3D object detection task. Our method achieved 76.79 mAP in moderate difficulty and outperformed many influential object detection models on the KITTI benchmark leaderboard.

Read the paper · More papers on PaperTik