Self-Supervised Multimodal Large Model for Cross-Language and Cross-Modal Knowledge Transfer

Chaofan Liao, Zexuan Jin, Zhiwei Zhang · 2025

Aiming at the alignment difficulties and migration bottlenecks of multimodal semantic modeling in cross-language and cross-modal scenarios, a large-scale multimodal representation model based on self-supervised learning is constructed, and a unified cross-language/cross-modal knowledge migration framework is proposed. The model combines image, text and audio multi-source data, models the shared semantic space through multi-task self-supervised pre-training mechanism, and introduces multiple contrastive learning and three-modal collaborative representation methods to effectively improve the inter-modal and inter-language semantic consistency. Systematic experiments are conducted in tasks such as graphic retrieval, cross-language Q&A, and multi-language description generation, and the results show that the model significantly outperforms mainstream contrastive methods in a number of evaluation metrics, and demonstrates excellent migration ability and generalization performance.

Read the paper · More papers on PaperTik