Noise robustness for HMM-based speech recognition systems
Adoram Erell · 2002
The problem is that of a mismatch in the level background noise between the training and recognition phases; the probability distributions estimated in the training phase are then no longer valid for the tested speech. Many different algorithms address this problem, being roughly classified into categories (1) augmenting the front end by a statistical estimator to estimate the clean speech parameters from the noisy signal; (2) adaptation of the HMM output PDs to the presence of noise; (3) modifying the front end so that the acoustic features are more robust to noise. The estimation approach has an inherent limitation: the information on the relative accuracy of different features does not get passed to the recognizer. A more rigorous probabilistic approach, applicable when the training speech database is clean relative to that in the recognition phase, is to adapt the probability computation to the noise instead of estimating the clean features.>