Speech enhancement using voice source models

Anisa Yasmin, Paul Fieguth, Deng Li · 1999

Autoregressive (AR) models have been shown to be effective models of the human vocal tract during voicing. However the most common model of speech for enhancement purposes, AR process excited by white noise, fails to capture the periodic nature of voiced speech. Speech synthesis researchers have long recognized this problem and have developed a variety of sophisticated excitation models, however these models have yet to make an impact in speech enhancement. We have chosen one of the most common excitation models, the four-parameter LF model of Fant, Liljencrants and Lin (1985), and applied it to the enhancement of individual voiced phonemes. Comparing the performance of the conventional white-noise-driven AR, an impulsive-driven AR, and AR based on the LF model shows that the LF model yields a substantial improvement, on the order of 1.3 dB.

Read the paper · More papers on PaperTik