A Multimodal Deep Network for Music Emotion Recognition Using Audio Chorus and Lyrics

Mohammad Ali Talaghat, Elham Parvinnia, Mahdi Mehrabi, Reza Boostani · IEEE Access · 2025

Music emotion recognition (MER) is an essential branch in music information retrieval, focusing on categorization of music based on emotional content. This study introduces a multimodal deep learning architecture adopting audio and lyrics to improve the emotion recognition accuracy. The model employs convolutional neural networks, long short-term memory layers, in addition with an attention to emphasize emotionally salient features. Significant performance gains are achieved by analyzing the chorus—often the most expressive and repetitive part of music. Herein, a new dataset containing 9,087 tracks which are labeled by valence and arousal, is prepared and introduced. Moreover, embeddings generated by XLNet and BERT networks are compared for lyrics feature extraction. Experimental results demonstrate the proposed scheme is superior to state-of-the-art methods in achieving significantly superior emotion recognition accuracy, which underscores the value of chorus-based analysis, deep attention networks, and multimodal integration in advancing MER.

Read the paper · More papers on PaperTik