Large Model Empowered Multi-Modal Semantic Communication With Selective Tokens for Training
Jincheng Peng, Huanlai Xing, Zhiwen Xiao, Lexi Xu, Xianfu Lei · IEEE Signal Processing Letters · 2025
Multi-modal semantic communication (MSC) has gained great attention due to its multi-modal processing ability. However, the existing MSC systems are mainly built on multi-modal large models that lead to inefficient computation on non-essential tokens, potentially restricting MSC from achieving more advanced levels of intelligence. To address this challenge, we propose a large model-empowered MSC system with a cross-modal attention-based token selection mechanism, denoted as LMECM-SC, which effectively utilizes the attention score across multi-modal tokens to filter out noisy or unuseful tokens, selectively learning the tokens that best benefit downstream applications. Meanwhile, we introduce the multi-modal adaptive semantic encoder and decoder that dynamically assign weights to encode multi-modal semantic information extracted from the selected tokens based on their modality and integrate semantic information with cross-modal attention scores at the receiver, optimizing the performance on downstream tasks. Experiment results indicate that LMECM-SC effectively reduces the number of tokens used for training, outperforming four baseline methods in terms of bilingual evaluation understudy score for text, learned perceptual image patch similarity for image, and perceptual evaluation of speech quality score for speech.