Improving naturalness in text-to-speech synthesis using natural glottal source
K. Matsui, Steve Pearson, K. Hata, Toshitaka Kamai · 1991
Various methods to improve text-to-speech in its naturalness and its ability to model individual speakers are discussed. Methods using a natural glottal source which is extracted from natural speech by an inverse-filtering technique are described. One method uses a repeating loop. Another method creates a source waveform of the desired pitch by concatenating single pulses. A multisource method which utilizes different types of glottal source by cross-fading techniques is proposed. Perceptual listening tests were performed with synthetic stimuli. The preliminary results show that these methods have the potential to improve the quality of text-to-speech synthesis.>