Synthesis and evaluation of natural sounding speech using the linear predictive analysis-synthesis scheme
Jayant M. Naik · University of Florida Digital Collections (University of Florida) · 1984
Linear Predictive coding is a popular technique for speech analysis and synthesis. But, continuous speech generated from a linear predictive (LP) synthesizer still lacks naturalness. The objective of this study is to examine the various issues involved in the production and evaluation of natural sounding synthetic speech, using the linear predictive model. In particular, a scheme for obtaining the control parameters of the LP synthesizer very reliably using the electroglottograph (EGG) signal as a glottal sensor is proposed. Our database consists of speech and EGG signals for sentences produced by male, female and child speakers. The two signals are obtained time synchronously. The features of the EGG signal and their relation to the glottal vibratory cycle are briefly discussed. Algorithms for computing the fundamental frequency contour and voiced/unvoiced decision from the EGG signal are outlined. Pitch synchronous LP analysis and synthesis schemes, guided by the EGG signal, are discussed. Excitation signals for synthesizing voiced speech are also derived from the EGG signal. The synthesized sentences are evaluated by a total of twenty listeners in a formal listening test. The results show that synthesis naturalness can be significantly improved with the use of the EGG signal. Errors in voicing/unvoicing and pitch computation are eliminated. Pitch synchronous analysis-synthesis performed over one whole period and non-impulse excitations derived from the EGG signal result in large improvements to synthesis naturalness.