Modeling the dynamics of speech and noise for speech feature enhancement in ASR

Stefan Windmann, Reinhold Haeb‐Umbach · IEEE International Conference on Acoustics Speech and Signal Processing · 2008

In this paper a switching linear dynamical model (SLDM) approach for speech feature enhancement is improved by employing more accurate models for the dynamics of speech and noise. The model of the clean speech feature trajectory is improved by augmenting the state vector to capture information derived from the delta features. Further a hidden noise state variable is introduced to obtain a more elaborated model for the noise dynamics. Approximate Bayesian inference in the SLDM is carried out by a bank of extended Kalman filters, whose outputs are combined according to the a posteriori probability of the individual state models. Experimental results on the AURORA2 database show improved recognition accuracy.

Read the paper · More papers on PaperTik