Incorporating Speaker’s Speech Rate Features for Improved Voice Cloning

Qin Zhe, Katunobu Itou · 2023

We investigate a neural network-based text-to-speech (TTS) synthesis system that aims to simulate the Mandarin voice of different speakers using short voice samples. Our system introduces new attempts in two key parts.: (1)Unlike prior studies, we not only capture the unique voice characteristics of each speaker but also integrate a speaker feature extraction model to improve our model’s effectiveness. Importantly, we utilize a duration model for accurately capturing the speech rate and rhythm of each individual, thus allowing our synthesized speech to demonstrate enhanced realism and credibility.(2) Based on the Tacotron model, we combine additional properties such as "speech rate" and "D-vector" to generate mel spectrograms from a given text, thereby enhancing the ability to clone through mechanisms such as location-sensitive attention. This ensures the quality and authenticity of the synthesized speech. Experiments have shown that this innovative approach uses less data to train and has better sound quality than previous models, and has a certain degree of reproduction in speech speed.

Read the paper · More papers on PaperTik