Modality Weights Based Fusion Model for Social Perception Prediction in Video, Audio, and Text
Eunchae Lim, Hyeon-Ji Yang, Hyung-Jeong Yang, Soo-Hyung Kim, Seungwon Kim, Ji-eun Shin, Aera Kim · 2024
Social perception is a crucial psychological concept that explains how we understand and interpret others and their behaviors. It encompasses the complex process of discerning individual characteristics, intentions, and emotions, significantly influencing social interactions and decision-making. In this paper, we propose a modality weights based fusion model for predicting 16 social attributes in the MuSe 2024 Challenge subtask MuSe-Perception. The proposed model utilizes visual, audio, and text data from the Chief Executive Officers' (CEOs') interviews to predict these 16 social attributes. This study mainly focuses on audio feature learning and the cross-attention of audio model outputs. Furthermore, we develop each modality weight-based fusion network that combines the outputs of each modality according to their weights. Experimental results show that the proposed model achieved competitive performance for some social attributes, but there were limitations to attain consistent overall performance. Based on these results, future work includes collecting more extended CEO video data and learning the importance of each modality for different attributes. This study is expected to contribute to developing a system for analyzing investment potential by predicting the CEOs' social attributes.