CDF-Net: A zero-shot 3D classification network with CLIP Decoder for Feature Fusion

Hao Yan, Jing Bai, Lu Liu · 2024

Zero-shot 3D model classification requires neural networks to correctly classify models with unseen categories, which has important practical implications for fields such as autonomous driving and robotics. PointCLIP demonstrates that transferring knowledge from the 2D image domain to the 3D model domain is beneficial for zero-shot 3D model classification, even without incorporating new 3D-specific knowledge. However, its effectiveness is limited by the lack of 3D knowledge. As a result, subsequent improvements have focused on refining 3D-to-2D projection algorithms rather than exploiting 3D information. Our approach aims to exploit both image domain knowledge and the full potential of 3D models. To this end, we propose CDF-Net, a learnable model architecture. CDF-Net consists of two components: an initial feature encoding module using the Contrastive Language-Image Pre-training network to extract 2D knowledge from multi-view representations of 3D models, and a learnable module to extract 3D knowledge and integrate it with 2D knowledge. Evaluations on two benchmarks consisting of four datasets show that CDF-Net achieves competitive classification accuracies, demonstrating the effectiveness of our method.

Read the paper · More papers on PaperTik