HybridBEV: Hybrid Encode and Distillation for Improved BEV 3D Object Detection

Junyin Wang, Chenghu Du, Huikai Liu, Zhenchang Xia, Bingyi Liu, Shengwu Xiong · IEEE Transactions on Intelligent Transportation Systems · 2025

The development of surround-view cameras is crucial for the advancement of autonomous driving. Utilizing depth information and image features to simulate LiDAR bird’s-eye-view (BEV) features can accomplish efficient 3D object detection tasks. Existing dense BEV generation methods heavily rely on the use of depth features, however, the suboptimal exploitation of these features often results in ambiguity in object location and feature representation during the BEV generation process. To address this, we have designed a hybrid encode and distillation method to enhance 3D object detection performance, termed HybridBEV. Initially, we designed the HybridEncode module, which employs a resampling strategy of depth features in voxel space to obtain BEV features that more accurately reflect the distribution of objects. Subsequently, we introduced multiple distillation methods to supervise the network’s voxel features and BEV feature representations, assisting the student network in learning critical features from the teacher model and ensuring that BEV features can more distinctly represent object distribution. Furthermore, during network training, we loaded pre-trained weights from the teacher network to guide network optimization and accelerate training. Extensive experiments on the nuScenes benchmark demonstrate that HybridBEV can effectively improve the performance of the student network and outperform previous state-of-the-art methods based on surround-view cameras. The code will be published athttps://github.com/wjyxx/HybridBEV

Read the paper · More papers on PaperTik