3D Point Cloud Pre-Training with Knowledge Distilled from 2D Images

Yuan Yao, Yuanhan Zhang, Zhenfei Yin, Jiebo Luo, Wanli Ouyang, Xiaoshui Huang · 2024

The success of pre-trained 2D vision models can largely be attributed to their ability to learn from large-scale datasets. However, compared with 2D image datasets, current pre-training data for 3D point clouds are limited. In this paper, we propose a knowledge distillation method for pre-training 3D point cloud models by directly acquiring knowledge from a 2D representation learning model, specifically the image encoder of CLIP. To close the significant domain gap between 2D images and 3D point clouds, we propose to align the features from the two domains at the concept level. Our method utilizes a cross-attention mechanism to extract concept features from 3D point clouds and compares them with the corresponding information from 2D images. This approach bridges the two domains and allows point cloud models to learn directly from the rich information contained in 2D teacher models. Extensive experiments show that our proposed knowledge distillation scheme achieves higher accuracy than the state-of-the-art 3D pre-training methods for synthetic and real-world datasets on various downstream tasks, including object classification, object detection, semantic segmentation and part segmentation.

Read the paper · More papers on PaperTik