G-RNN-GAN for Singing Voice Separation
Hui Zhang, Niannao Xiao, Peishun Liu, Zhicheng Wang, Ruichun Tang · 2020
Singing voice separation is required for many applications such as interactive music players, MIR (Music Information Retrieval), etc. Mixed music including both background music and human voice which can be used in different areas. We propose a network combining RNN and GAN to improve the separation performance of singing voice, named G-RNN-GAN. The importance of RNN used in the generator also be demonstrated. By generating human singing voices from mixed music in the generator, we discriminate the output in discriminator. The loss function is defined as the sum of voice adversarial loss and conditional voice loss. The latter guarantees that the network can be fitted properly. Experimental results show that our network can converge in less than 30 minutes. Experiments on MIR-1K data set show that the values of GNSDR and GSAR reach 8.232 and 9.605 and GSIR has a high value of 14.798.