Multimodal Affect Perception With Large Language Model Enhancement Network

Kaixiang Yang, Yifan Luo, Zongyan Zhang, C. L. Philip Chen, Tong Zhang · IEEE Transactions on Affective Computing · 2025

Multimodal Sentiment Analysis (MSA) plays a vital role in understanding emotional content from social media and multimedia data. However, existing methods often rely on large-scale labeled datasets, leading to high annotation costs and poor adaptability. They also suffer from modality imbalance and suboptimal feature fusion. To address these issues, we propose MapleNet—a Multimodal Affect Perception framework enhanced by Large Language Models. MapleNet integrates a prototype-guided fusion strategy and a dynamic modality balancing mechanism to improve alignment and collaboration between text and image features. Specifically, a shared-space encoder combined with prompt optimization ensures semantic consistency across modalities. Within the prototype learning framework, the model dynamically adjusts modalityspecific learning by aligning features with class prototypes, thus mitigating imbalance and uncovering complementary affective cues. In addition, MapleNet employs a similaritybased sample retrieval module to construct contextual prompts, enriching sentiment understanding in few-shot settings. Experiments on six benchmark datasets show that MapleNet consistently outperforms state-of-the-art methods, especially under few-shot conditions, achieving superior accuracy and generalization. The relevant code is available at https://github.com/YFanLuo/MapleNet.

Read the paper · More papers on PaperTik