A context adaptation approach for building context dependent models in LVCSR

Xiaoxing Liu, Baosheng Yuan, Yonghong Yan · 2001

Abstract This paper introduces a new context adaptationframework for building context dependent HMM modelsin LVCSR. In this new framework, all states of eachcenter phone are clustered into groups by the decisiontree algorithm. All the tied states of context dependentHMM models were then derived by adapting theparameters of the multiple-mixture context independentmodel via data dependent MAP (maximum a posterioriprobability)method using the training vectorscorresponding to the tied state. An advantage of thisapproach is that it can maintain a high prediction andclassification power given limited training data thereforethe model trained in this framework is more reliable thanin conventional framework. Experimental results onWall Street Journal corpora demonstrate that theproposed approach leads to a significant improvement inrecognition performance. 1. Introduction Decision tree state tying based context modeling hasbecome increasingly popular for modeling speechvariations in large vocabulary speech recognition[1][2].In the conventional framework the stochastic classifierfor each tied state is trained using Baum-Welchalgorithm using the training data corresponding to thespecific tied state[3]. However, the context dependentclassifiers trained using this method are not so reliablefor the training data corresponding to each tied state islimited and model parameters are easily to be affectedby undesired sources of information such as speaker andchannel differences contained in the training data. Toattack this problem, We propose a new contextadaptation method to estimate the parameters of contextdependent models. In this method, a multiple-mixturecontext independent model is trained firstly. In decisiontree clustering, single mixture Gaussian models wereused to establish the state tying. After all the contextdependent states are clustered into groups, the clusteredstates of context dependent model were derived byadapting the parameters of the context independentmodel via data dependent MAP method using trainingdata corresponding to the tied state. We consider themulti-mixture context independent model as coveringthe space of more broad speaker and environmentclasses of speech signal, then adaptation is the contextdependent tuning of those speaker and environmentclasses observed in context’s training speech. Mixtureparameters for those speaker and environment classesnot observed in the training speech of the specific tiedstate are merely copied from the context independentmodel. This means that the model has higher predictionand classification power for the test data from speakerand environment classes unseen or rarely seen in thecontext’s training data. Experimental results on WallStreet Journal corpora demonstrate that the proposedapproaches lead to a significant improvement inrecognition performance.

Read the paper · More papers on PaperTik