A glottal waveform model for high-quality speech synthesis

Seiichi Tenpaku, Tatsuya Hirahara · The Journal of the Acoustical Society of America · 1990

A new glottal waveform model for high-quality speech synthesis is proposed and the results of the perceptual evaluations for synthesized speech using the proposed model and other models are compared. The proposed glottal waveform model consists of two parts: a waveform generator and a spectrum shaping filter. A third-order polynomial, of which coefficients are determined by combinations of OQ (open quotient), SQ (speed quotient), AV (amplitude of voicing), and F0, is used for the waveform generator. A second-order IIR filter, which is designed to control the spectral tilt and the relative amplitudes of lower harmonics components by two parameters, serves as the spectrum shaping filter. Thus the parameters have a direct effect on the waveform and its spectral shape. Using the three kinds of information (F0, power, and formant) extracted from the eight different Japanese words pronounced by a male and a female announcer, 80 synthesized speech stimuli were prepared for the preference test. The stimuli were generated by a cascade formant synthesizer using five different glottal waveform models: the proposed model, Fant's model [STL-QPSR 4, 1–13 (1985)], Fujisaki's model [ICASSP 86, 1605–1608 (1986)], Klatt's model [J. Acoust. Soc. Am. 67, 971–995 (1980)], and Rosenberg's model [J. Acoust. Soc. Am. 49, 583–590 (1971)]. Results of the preference tests with 20 subjects show that the naturalness of the synthesized speech generated by the proposed model is as good as those of the Fant and Fujisaki models.

Read the paper · More papers on PaperTik