Comparative Study on Multi-voice Singing Synthesize Systems
S. Resna, Rajeev Rajan · International Journal of Automation and Smart Technology · 2023
In this paper, two multi-voice singing synthesis frameworks are compared One proposed model consists of two blocks, namely, text-to-speech (TTS) converter and speech-to-singing (STS) converter. Synthesized speech is generated from lyrics for a target speaker's voice by TTS converter in the front-end. Later, a sung version is synthesized as per the given target-melody using encoder-decoder model in the STS module. We have compared our model with an existing multi-voice singing synthesis model, based on generative adversarial network (GAN) with phoneme synchronization information. The proposed system is systematically evaluated using subjective and objective tests. Three performance metrics, namely the mean opinion score (MOS), log spectral distance (LSD) have been analyzed as part of the study. Our study shows that the proposed model generates singing voices that adapts well to the target melody but the phonetic intelligibility is poor when compared to the baseline system.