CAUD-3D: Quantifying Cross-Modal Alignment and Uni-Modal Disentanglement for 3D Understanding via Distribution Similarity Coefficient
Yuanlei Hou, Zongqi Liu, Jiajing Hu, Linlin Yang · 2025
The lack of suitable evaluation metrics hinders the precise measurement of biases in the cross-modal feature space and the distinctiveness of 3D point cloud features, impeding further optimization efforts for enhanced 3D understanding. To tackle these challenges, we present the unified distribution similarity coefficient driven multimodal pre-training for 3D understanding framework, termed CAUD- 3D, providing a deeper understanding of the cross-modal alignment and uni-modal disentanglement process of the multimodal pre-training. Specifically, we generalize class-wise features to a Gaussian distribution, facilitating the quantification of representation quality within the hyper-sphere space through the calculation of the distribution similarity coefficient. To the best of our knowledge, this is the first work to measure the representation quality of cross-modal features from the perspective of the distribution similarity coefficient. Furthermore, we formulate the cross-modal class-wise alignment and uni-modal class-wise discrepancy loss terms to align cross-modal class-wise feature distribution and disentangle the interference among the 3D feature class-wise distribution. Our method significantly outperforms previous works.