Mandarin Singing Synthesis Based on Generative Adversarial Network
Yun Hong Zhou, Hongwu Yang, Ziyan Chen, Yajing Yan · 2020
This paper proposed a method for statistical parametric singing synthesis incorporating GAN (Generative Adversarial Network) that trained acoustic model. In GAN, the acoustic model was trained to minimize the weighted sum of the conventional minimum generation loss and adversarial loss, which was minimizing the distance between the natural and generated samples parameter, thus effectively solved the problem of over-smoothing. In the experimental part, we established a singing voice corpus with 60 songs and divided them that have been recorded and labeled into about 1000 sentences, of which 950 sentences were for training model. Comparing the generated songs of the method proposed in this paper and HMM, through 10 people MOS scores, the score of the former was 3.12 that was better than the latter of 2.81.