Multimodal Multifaceted Music Emotion Recognition Based on Self-Attentive Fusion of Psychology-Inspired Symbolic and Acoustic Features
Jiahao Zhao, Kazuyoshi Yoshii · 2023
This paper describes automatic music emotion recognition (MER) that aims to estimate the valence and arousal (V/A) scores of a piece of piano music. The emotion is multifaceted in nature; it is rendered by various features that are often mutually dependent and inherent in music composition and performance. A basic approach to MER is to train a deep neural network (DNN) that extracts latent features representing the emotion as a whole and estimates the V/A scores, using only a limited amount of audio data with imbalanced V/A annotations. Such a black-box approach, however, suffers from limited performance and interpretability. To overcome these limitations, in this paper we propose a multimodal multifaceted MER method that fuses various kinds of musically-meaningful symbolic and acoustic features extracted from both MIDI and audio data, respectively, based on the expert knowledge of musical psychology. More specifically, our method separately extracts the affective features representing the rhythm, dynamics, melody, harmony, and tone color of a piano piece as the main factors affecting the emotion and integrates them with a self-attention mechanism that can learn the complicated cross-modal relationships. The experiments using the common EMOPIA dataset showed that the proposed model achieved the state-of-the-art V/A classification accuracy of 69.2% and that the multimodal and multifaceted feature fusions contributed to the performance improvement.