LLM-Based Expressive Text-to-Speech Synthesizer with Style and Timbre Disentanglement
Yuanyuan Zhu, Jiaxu He, Ruihao Jing, Yaodong Song, Jie Lian, Xiao-Lei Zhang, Jie Li · 2024
The ICACG challenge Track 1 requires to generate the target speaker audios with high naturalness under extremely limited speaker dataset. To achieve this goal, we introduce a novel text-to-speech synthesizer which disentangles the style and tim-bre information in cascade approach. Firstly, an auto-regressive large language model (LLM) is applied to complete the text-to-token generation, which can capture the style information from audio prompt in zero-shot mode. Subsequently, a variational generator is used to reconstruct the mel-spectrogram in corresponding to the target speaker timbre conditioned on speaker embedding. Therefore, the final synthesized audio can not only contain the timbre of the target speaker, but also achieve a high degree of expressiveness by utilizing the capabilities of LLM. Since large and diverse data is necessary for training, a novel data processing pipeline is also proposed to process the collected data. As a result, our system achieved great performance in terms of expressive speech synthesis and ranked the first place in ICAGC 2024 Track 1 over all five evaluation metrics: 3.89 Quality, 3.83 Similarity, 3.85 Emotion, 3.89 MOS(avg) and 0.22 MOS (std).