Construction of multimodal music automatic annotation model based on neural network algorithm

Zhong Miao, Chaozhi Cheng · 2023

Improving the effect of music annotation task through various advanced technologies has become a hot direction in the field of music information retrieval. The research of music automatic annotation has high value in both theoretical research and practical application. The deep CNN (Convolutional Neural Network) model in the field of deep learning has achieved good results in the fields of image and voice. In this paper, the construction of multimodal music automatic labeling model based on neural network algorithm is launched. In this paper, CNN combined with SAM (Self-Attention Mechanism) is used to learn the appropriate feature representation from the low-level Mel spectrum description of music and the original audio waveform data. Two-dimensional convolution is applied to the Mel spectrum input of music, and one-dimensional convolution is applied to the original audio waveform input, so as to better capture various structural features of music and annotate it. The results show that the accuracy of CNN combined with SAM method is 3.9% higher than that of linear weighted fusion method, and the AUC value is 8.5% higher than that of linear weighted fusion method. The comparison results show that the multi-modal music automatic annotation model framework proposed by CNN and SAM in this paper is effective for the automatic annotation task of music.

Read the paper · More papers on PaperTik