Usable speech processing: A novel approach to processing speech in degraded environments

Brett Y. Smolenski · The Journal of the Acoustical Society of America · 2004

One of the main challenges still plaguing speech processing applications is enabling them to work in operational environments where interference and noise abound. The traditional approach has been to use some form of adaptive filtering operation. However, since the speech is nonstationary, it is possible to extract segments from the speech signal that have a large segmental signal-to-noise ratio (SNR) even when the overall SNR is very low. Such high SNR segments frequently occur during voiced speech, and experiments have shown that, using a speaker identification system, these high SNR segments can be correctly identified even when the entire utterance cannot. However, the segmental SNR is not normally known a priori. In this research, a statistical model is first developed for the segmental SNR values for several commonly occurring environments. It is then shown that a modified sinusoidal model and the Teager energy operator can be combined using a context dependent nonlinear regression technique to obtain a low variance estimate of the segmental SNR values provided the frame size is larger than 30 samples.

Read the paper · More papers on PaperTik