Speech Synthesis Method Based on Tacotron2

Li Yang, Donghong Qin, Jinbo Zhang · 2021

Compared with traditional speech synthesis systems, end-to-end speech synthesis systems based on deep learning (such as DeepVoice3, Tacotron2) not only reduce the requirements for linguistic knowledge, but the synthesis effect is almost close to the level of human pronunciation. However, the end-to-end speech synthesis system based on deep learning has disadvantages such as missing words, repeated pronunciation, and slow synthesis speed. In view of the local information preference of the Tacotron2 model in the decoder, this paper proposes to maximize the interactive information between the text and the predicted acoustic features and use the WaveGlow synthesizer to reduce the local information preference and the problem of slow synthesis speed, pronunciation in the Tacotron2 model. Experimental results show that the improved model subjective evaluation MOS (Mean Opinion Score) score is 3.94, and the synthesis speed is significantly improved.

Read the paper · More papers on PaperTik