Non-negative Matrix Factorization and DenseNet based Monophonic Singing Voice Separation

Kun Zhao, Zhao Zhang, Zheng Wang, Yutian Wang · 2022 IEEE 5th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC) · 2022

Most of traditional singing voice separation methods usually assume that the vocal model is the source-filter model thatextracts the transfering functions of the vibration source and the filters separately. In recent years, with the rapid development of deep learning, end-to-end methods have become increasingly popular and achieved better separation results. Deep neural networks are very useful for processing complex nonlinear data. However, such models usually have large parameter sizes and lack interpretability. In this paper, we propose a novel singing voice separation model which combine the advantages of both source-filter model and deep neural network. We use DenseNet to extract the fundamental frequency of singing voice by using its powerful feature extraction ability, and use human voice resonance model and non-negative matrix factorization (NMF) to obtain the soft mask to get the final separated voice. Experimental results show that our system can obtain better separation effect and generalization performance.

Read the paper · More papers on PaperTik