FGP-GAN: Fine-Grained Perception Integrated Generative Adversarial Network for Expressive Mandarin Singing Voice Synthesis

Xin Liu, Weiwei Zhang, Zhaohui Zheng, Mingyang Pan, Rong Wang · IEEE Transactions on Consumer Electronics · 2024

Though singing voice synthesis (SVS) has been explored by recent works, pitch over-smoothing and spectral blurring are still unresolved, resulting in lack of expressiveness of the singing voice. To tackle the afore-mentioned problems, an SVS model that adopts fine-grained perception modeling in the generative adversarial framework is proposed in this work. Specifically, the forward difference of fundamental frequency is also modeled in addition to the conventional fundamental frequency in the generator, resulting in more accurate vocal fundamental frequencies. Moreover, the fine-grained perception module is utilized to restrict the generator to focus more on detailed information. Then, we adopt the WGAN discriminator to optimize the SVS model such that the distribution of the generated singing spectrogram effectively approximates the actual one. Finally, the spectrogram was fed into our modified vocoder, resulting in more expressive and natural singing voice. The effectiveness of the proposed approach is verified on a professional Mandarin Chinese corpus. Experimental results demonstrate that the proposed approach can obtain more accurate fundamental frequencies, clearer spectrograms, more natural singing voice and higher mean opinion score (MOS), comparing with several state-of-the-art approaches.

Read the paper · More papers on PaperTik