Hybrid time- and frequency-domain speech synthesis with extended glottal source generation
Georg Fries · 2002
A novel synthesis approach that combines speech synthesis in the time domain with speech synthesis in the frequency domain is introduced. The intention is to improve speech quality by designing a hybrid system which profits from the advantages of both methods and overcomes some of their drawbacks. Compared to a stand-alone formant synthesizer, a better quality of fricatives and plosives has been achieved, whereas the flexibility in fundamental frequency variation is preserved. Moreover, simultaneous use of both system components enables the system to produce naturally sounding transitions at the segment boundaries. The parametric part of the hybrid system-a formant-based synthesizer-is excited with a time-domain source generation scheme. It is based on concatenation and modification of stored natural source waveforms. Important system features are phoneme-specific variants of stored source waveforms and additional generation of shimmer and jitter. Preliminary informal listening tests showed that the naturalness of the voiced sounds has been improved compared to the results of the previous synthesizer.>