Domain Adaptation on Point Clouds via Multi-Modal Representation Learning

Longkun Mao, De Xing, Xiaorong Zhang, Guijuan Wang · 2024

Enhancing cross-domain performance of 3D point cloud neural networks remains a formidable task due to subtle variations in feature distributions among datasets, limiting their generalizability beyond the training domain. To tackle this, we have introduced an innovative approach that capitalizes on the strengths of 2D and 3D modalities, specifically aimed at bolstering unsupervised domain adaptation (UDA) in object point cloud classification. Our method entails dual branches: one transforms 3D point clouds into 2D depth maps, employing the visual encoder of CLIP [1] to convert these maps into 2D feature representations, while the other processes point cloud data through a 3D representation network for comprehensive 3D feature extraction. The crux of our innovation lies in a tailored adapter meticulously harmonizing information from these diverse feature modalities, augmenting the network's perceptual abilities across both the source and target domains. Empirical evaluations on the PointDA-10 dataset validate our approach's effectiveness, showcasing its promise in unsupervised domain adaptation and presenting a fresh perspective for addressing cross-domain challenaes in the realm of point clouds.

Read the paper · More papers on PaperTik