HMM-based Thai Singing Voice Synthesis System
Lattapon Jeerapradit, Atiwong Suchato, Proadpran Punyabukkana · 2018
Singing synthesis in each language has its unique characteristics and challenges aiming to improve its naturalness. One of the challenges that tonal language singing synthesis specifically faces is melisma, which is a syllable with more than two musical notes. This research offers a melisma-compatible singing voice synthesis system. To do so, we propose phoneme duplication methods. The results showed that the proposed phoneme duplication methods made the system compatible with melisma, where short vowels and final consonants constructed a favorable waveform closer to actual singing voice. It also provides higher naturalness as perceived through MOS evaluation. Finally, in order to improve naturalness in the synthesized singing voice, an experiment with HMM state numbers was conducted. The outcome demonstrated that the naturalness increased as the state numbers grew to a certain point.