HMM Adaptation Using Statistical Linear Approximation for Robust Speech Recognition
Michael Berkovitch, Shallom D.Il · InTech eBooks · 2011
Automatic Speech Recognition (ASR) systems, show degraded recognition performance when train and operate on mismatched environments.This mismatch can be caused due to different microphones, noise conditions, communication channels, acoustical environment etc.This work is motivated, in part, by the Distributed Speech Recognition (DSR) architecture.The DSR uses ASR server that provides speech recognition services to different devices that may operate in different environments (i.e.mobile devices).Thus, the ASR server must implement environment compensation techniques.The traditional ASR environment compensation techniques use filtering and noise masking, spectral subtraction and multi microphones array.These techniques are usually implemented in the ASR front end and aims to provide clean speech samples to the ASR engine.State of the art ASR systems use Hidden Markov Models (HMM) to represent the stochastic nature of the speech features.These statistical models achieves high recognition rate when trained and tested at the same environmental condition.To add noise robustness for these models, methods such as Maximum Likelihood Linear Regression (MLLR) and Parallel Model Combination (PMC) had been developed.These methods perform an adaptation of the ASR engine to better fit the recognition environment.The main drawback of these methods is there computational complexity and the need of large adaptation data, which makes them not suitable for real-time application.The environment compensation technique, presented in this chapter, is an extension of the Statistical Linear Approximation (SLA) method originally applied in the feature space to the model space.Using this environment compensation technique, new adapted HMMs set are created Using the clean speech HMMs and the noise model.The adapted HMMs are then used for the recognition of the degraded speech.The proposed robustness method is highly attractive for the Distributed Speech Recognition (DSR) architecture, since there is no impact on the Front End structure and neither on the ASR topology.Experiments, using this method, show high recognition rates in various noise conditions, close to the case of matched training (i.e.recognition and training performed in the same degraded environment). Technique for Robust Speech RecognitionTechniques for Robust Speech Recognition can be divided into three categories.www.intechopen.com