Iterative self-learning speaker and channel adaptation under various initial conditions

Yunxin Zhao · 2002

A self-learning adaptation technique is presented which handles the speaker and channel induced spectral variations without enrolment speech. At the acoustic level, the distortion spectral bias is estimated in two steps using the unsupervised maximum likelihood estimation: in the first step, the probability distributions of the speech spectral features are assumed uniform for severely mismatched channels; in the second step, the spectral bias is reestimated assuming Gaussian distributions for the spectral features. At the phone unit level, unsupervised sequential adaptation is performed via Bayesian estimation from the online, bias-removed speech data, and iterative adaptation is further performed for dictation applications. Over four 198-sentence test sets, on a continuous speech recognition task with vocabulary size=853 and grammar perplexity=105, the largest increase of average word accuracy is 85.2% from the baseline accuracy of -0.3%, and the maximum average word accuracy is 89.4% from the baseline accuracy of 56.5%.

Read the paper · More papers on PaperTik