Multimodal Social Relationship Recognition Based on LLM
Haopeng Wang, Zhitian Zhang, Menglei Xia, Dejiao Huang, Ruyi Chang, Shuai Guo · IEEE Access · 2025
In recent years, multimodal social relation recognition has become a critical task in the fields of computer vision and natural language processing. However, existing research still faces key gaps, particularly in effectively aligning image features with linguistic features to improve recognition accuracy. This paper proposes a social relation recognition method based on multimodal feature fusion and validates the crucial role of cross-modal alignment mechanisms in enhancing recognition accuracy. Our approach leverages large language models to meticulously extract event structures from textual descriptions, capturing key elements such as emotional states, scenarios, and relationships, while employing convolutional neural networks to extract deep features from images. Subsequently, we introduce a cross-modal alignment mechanism to semantically align textual event structures with visual features, ensuring high semantic consistency between the two modalities. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms existing unimodal and basic multimodal approaches, confirming its effectiveness and innovative contributions to social relation recognition.