Topic-based language models using EM

Daniel Gildea, Thomas Frank Hofmann · 1999

The performance of global affine and nonlinear transformations for speaker adaptation in a hidden Markov model (HMM) speech recognition system are compared in this paper. The nonlinear transformation was obtained with a multilayer perceptron network (MLP) which was trained during the adaptation process to transform the mean vectors of the HMMs such that the output probabilities of the HMMs for the adaptation utterances were maximized. The performance of the MLP adaptation method was compared to the maximum likelihood linear regression (MLLR) adaptation procedure. Both of these methods were tested in a connected digit speech recognition system using multi-environment models. The results show that the nonlinear MLP transformation clearly outperforms MLLR in terms of adaptation speed. Moreover, the performance of MLP adaptation with larger amounts of data was comparable to the MLLR performance.

Read the paper · More papers on PaperTik