Fractal modeling of a glottal waveform for high-quality speech synthesis

Naofumi Aoki, Tohru Ifukube · The Journal of the Acoustical Society of America · 1998

From several psychoacoustic experiments, it is well known that the oversimplified glottal waveform degrades the quality of synthesized speech. For the enhancement of the quality of synthesized speech, it is desired to generate a glottal waveform in detail using an appropriate model which requires a few control parameters. In this study it is supposed that the differential glottal waveform, which is estimated by linear prediction filter, is modeled as a random fractal signal. The differential glottal waveform is regarded as the first integral of white noise whose frequency characteristic is approximated as −6 dB/oct. This signal can be interpreted as Brownian motion in terms of fractal theory. In order to generate random fractal signal, a kind of wavelet synthesis method called the midpoint displacement method is employed. Psychoacoustic experiments were conducted to examine whether or not the proposed method is effective. Experimental results indicate that the quality of original speech is fully preserved even if more than 90% of wavelet coefficients are replaced with a random fractal signal generated by the proposed method. The wavelet coefficients that cannot be replaced are to be the control parameters in this model.

Read the paper · More papers on PaperTik