Multi-View Text Enhancement for Parameter-Free Zero-Shot 3D Model Classification
Shuting Xi, Jing Bai, Zheng Hu · 2025
Large-scale pre-trained models have demonstrated significant advantages when dealing with visual and language tasks in open-world scenes. However, recent studies have identified certain limitations when utilizing comparative language-visual pretrained models for zero-shot 3D model classification, mainly in the form of neglecting the multi-view independent information of the 3D model. In addition, existing textual cues usually rely only on category semantics and fail to fully exploit the rich contextual structure of the 3D model itself, thus limiting the effective matching of visual and semantic information. To this end, this paper proposes a multi-view text-enhanced zero-shot 3D model classification based on a parameter-free network. The method utilizes a pre-trained image encoder to extract multi-view visual features and combines the semantic comprehension capability of a large language model to mine 3D model structure information from multiple perspectives. Fine-grained associations between views and textual cues are made through view-by-view interaction, and decision-level fusion is used to integrate the independent decision results of each view to achieve zero-shot classification. Without any 3D training, the method in this paper achieves 67.4% accuracy on the ZS3D dataset and 61.3%, 43.3% and 28.1% accuracy on the three sub-datasets of the Ali dataset, which verifies the generalizability and effectiveness of the proposed method.