Speech Synthesis : Improving Production Quality
Neal B. Pinto, D. G. Childers, A.L. Lalwani · 1989
We describe analysis and synthesis methods for improving the quality of speech produced by Klatt's software formant synthesizer (381. Synthetic speech, generated using an excitation waveform resem- bling the glottal volume-velocity, was found to be perceptually pre- ferred over speech synthesized using other types of excitation. In ad- dition, listeners ranked speech tokens synthesized with an excitation waveform that simulated the effects of source-tract interaction higher in naturalness than tokens synthesized without such interaction. A series of algorithms for silent and voiced/unvoiced/mixed excita- tion interval classification, pitch detection, formant estimation, and formant tracking were developed. These algorithms can utilize two channels of input data, i.e., the speech and the electroglottographic (EGG) signals, and can, therefore, surpass the performance of single channel (acoustic-signal-based) algorithms. The formant synthesizer was used to study some aspects of the acoustic correlates of voice quality, e.g., male/female voice conversion and the simulation of breathiness, roughness, and vocal fry.