Multimodal Audio Emotion Recognition with Graph-based Consensus Pseudolabeling
Gabriel Natal Coutinho, Artur de Vlieger Lima, Juliano Yugoshi, Marcelo Isaias de Moraes, Marcos Paulo Silva Gôlo, Ricardo Marcondes Marcacini · 2023
This paper presents a novel method called Multimodal Graph-based Consensus Pseudolabeling (MGCP) for unsupervised emotion recognition in audio. The goal is to determine the emotion of audio segments using the circumplex model of emotions. The method combines pre-trained unimodal models for audio and text and follows a three-step process. First, audio segments are represented using embeddings from unimodal models. Then, modality-specific graphs are constructed based on similarity and integrated into a multimodal graph. Finally, pseudolabels are generated by measuring consensus between modalities, and a graph regularization framework is introduced to estimate the final emotion coordinates. Experimental evaluation shows the effectiveness of the MGCP method, surpassing both unimodal and traditional multimodal models, enabling audio emotion recognition without labeled data specific to the target domain.