Noisy CMLLR for noise-robust speech recognition

Do Kyun Kim, Mark Gales · Cambridge University Engineering Department Publications Database · 2009

Adaptive training is a widely used technique for building speech recognition systems on non-homogeneous training data. Recently there has been interest in applying these approaches for situations where there is significant levels of background noise. Various schemes for adaptive training are based on noise, or speaker, specific transforms of the observed noise-corrupted speech to yield estimates of the clean speech. However when there are high levels of background noise, these clean speech estimates may be poor resulting in degradations in performance. In this work, a new approach for adaptive training on noise-corrupted training data is presented. It extends a popular form of linear transform for model-based adaptation and adaptive training, constrained MLLR (CMLLR), to reflect additional uncertainty from noise-corrupted observations. This new form of transform is called noisy CMLLR (NCM-LLR). NCMLLR uses a modified version of generative model between clean speech and noisy observation, similar to factor analysis (FA). However in contrast in FA here the generative model describes a transformation, rather than a covariance matrix structure. The use of NCMLLR for adaptation and adaptive training using an expectation-maximisation approach is described. Discriminative adaptive training with NCMLLR is also presented based on the minimum phone error criterion. Experiments are conducted on noise-corrupted version of Resource Management and in-car recorded digit data. In preliminary experiments this new approach achieves improvements in recognition performance over the standard approach in low signal-to-noise ratio conditions. In addition the need for adaptive training when there are a range of noise conditions in the training data is shown. 2 1

Read the paper · More papers on PaperTik