ConVoxel: From complementarity to consistency in multi-voxel 3D object detection

Ao Lei, Dejene Mengistu Sime, Nan Ouyang, Kai Sheng, Xiaojiang Ren, Yiting Liu, Xin He, Yoshiki Yamaguchi · Knowledge-Based Systems · 2026

3D object detection is a cornerstone technology for environmental perception in autonomous driving systems. However, existing voxel-based methods face key challenges. Single-resolution paradigms suffer from quantization artifacts. Conversely, multi-resolution paradigms that pursue feature complementarity by fusing disparate scales often introduce semantic conflicts and feature degradation. This paper challenges the conventional complementarity-based approach. We propose ConVoxel, a novel framework built on the principle of feature consistency. Instead of fusing heterogeneous features, ConVoxel generates multiple ‘views’ of the point cloud using proximate-resolution grids. These views are treated as redundant representations for verification. We first use a Shared Sparse Multi-view (SSM) backbone to extract features in a unified semantic space. Then, our novel Multi-view Feature Fusion (MVFF) network dynamically verifies and reinforces consistent object features while suppressing quantization noise. Extensive experiments on the authoritative KITTI dataset demonstrate that ConVoxel improves upon the SECOND baseline by 2.01% mAP, demonstrating the superiority of the consistency-based paradigm.

Read the paper · More papers on PaperTik