Research on Multi-modal Music Emotion Classification Based on Audio and Lyirc
Gaojun Liu, Zhiyuan Tan · 2020 IEEE 4th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) · 2020
To solve the problem of low accuracy of emotion classification, this paper proposes a new multi-modal fusion emotion classification method based on audio and lyrics. Firstly, Mel Frequency Cepstrum Coefficient, spectrum centroid and frequency-band energy distribution are used as feature data in audio, and LSTM in deep learning is applied to music emotion classification; In terms of lyrics, the Bert model is used to classify the lyrics, and the sentiment dictionary is used to perform LFSM-based equalization on the lyrics emotion classification results. Finally, a new fusion method is proposed on the traditional fusion method. The experimental results show that the new fusion method has 5.77% and 4.03% improvement over the linear weighted multimodal fusion and LFSM fusion methods.