PVS: a 3D object detection network based on point-voxel feature fusion and sequence prediction

Jun Yang, Erdun Zhao, Juan Yao · 2025

3D object detection is a key component of autonomous driving systems. However, the sparse and unstructured nature of point cloud data poses great challenges for feature extraction. In this paper, we propose a two-stage detection framework named PVS Net, which improves the detection performance by innovatively fusing point features and voxel features. In the first stage, the framework uses voxel features to capture global information, generating high-quality target proposals through sequence prediction. In the second stage, the local geometric details of point features are fused to refine the detection results. The approach focuses on enhancing both the stability and accuracy of detection. Specifically, a point-voxel feature fusion architecture is designed to improve feature representation, a robust ground plane estimation algorithm is proposed to optimize the data augmentation process, and an improved bounding box regression loss function is introduced to boost localization accuracy. PVS Net achieves a 68.14% mAP on the ONCE dataset, setting a new state-of-the- art, particularly in the detection of long-range targets (50m to infinity). This method demonstrates significant advantages for point cloud-based 3D object detection and offers new solutions to the challenges in this field.

Read the paper · More papers on PaperTik