State dependent feature component selection for noise robust ASR

Bert Cranen, J.M. de Veth · Radboud Repository (Radboud University) · 2004

The acoustic environment in w hich speech is recorded has a strong influence on the statistical distributions o f observed acoustic features.In order to make A S R in sensitive to noise it is crucial that these distributions are sim ilar in the training and testing condition.M ostly, it is attempted to compensate for the impact o f noise by estimating the noise characteristics from the signal.In this paper we explore the feasibility o f a new method to increase noise robustness: W e try to exploit a priori knowledge stored in clean speech models.U sing M e l bank log-energy features, recognition is done b y ignor ing the model components for features that contained lit tle energy during training.This strategy aims at recog nition results that are determined more strongly b y the match in the high-energy rather than b y the mismatch in the low-energy model components.Application o f the new method to clean speech data confirms that discard ing components below a certain energy threshold does not deteriorate recognition performance.Experiments with noisy data, however, show that performance gains are rel atively small.This paper explains w hy that is the case and w hy, despite the lim ited success, the outcomes suggest that the method still could prove to be a valuable addition to data-driven methods like (bounded) marginalisation.

Read the paper · More papers on PaperTik