CE-Tacotron2: End-to-End Emotional Speech Synthesis
Wang Zhi, Yinhua Liu, Liang Shan · 2021
Text-to-speech synthesis is an end-to-end synthesis technology, which can produce human-like synthesized speech through a computer. In the end-to-end speech synthesis system, Tacotron2 is an advanced speech synthesis model that can synthesize high-quality speech. However, it is quite different from the real speech since the emotion is not rich enough and the naturalness is insufficient. Therefore, an emotional speech synthesis model is proposed based on the Tacotron2 model for Chinese text-to-speech synthesis through embedding Chinese Preprocessing Module and Emotion Embedding Module. Experimental results show that the model successfully expresses various emotional states into synthesized speech. The subjective listening test evaluates the naturalness and emotional expression of synthesized speech, which verifies the feasibility of the model.