LLM-Driven Multimodal Fusion for Human Perception Analysis
Sergio Esteban-Romero, Iván Martín-Fernández, Manuel Gil-Martín, David Griol, Zoraida Callejas, Fernando Fernández-Martínez · 2024
The Multimodal Sentiment Analysis Challenge presents two distinct sub-challenges related to human perception characteristics. This paper focuses on the MUSE-PERCEPTION challenge, which aims to predict the perceptual attributes of CEOs from video data. Specifically, we propose a novel approach that leverages Multimodal Large Language Models (MM-LLM) to integrate multimodal features and generate accurate predictions. Our proposal explores both the direct generation of target values using tuned Qwen-VL and the extraction of semantic representations from a pre-trained Gemma-2B model. The latter integrates multimodal data such as facial expressions (vit-fer), audio (w2v-msp) and additional features such as egemaps or facenet512 results. The input data involve a task-specific cue and an audio transcript. Feature projectors map each modality into the semantic space of the LLM, and mean or attention clustering aggregates the final hidden state of the LLM. Next, an MLP uses these aggregated representations to predict 16 perceptual attributes. To optimize model performance, we experimented with several fitting methods and implemented a combined mean absolute error (MAE) and Pearson correlation coefficient (ρ) loss function. The experimental results shows the effectiveness of this approach, and the Gemma-2B-based model achieves a solid second place in the challenge. These results underscore the potential of MM-LLMs to improve the understanding of human perception in professional contexts.