A speaker adaptation technique using linear regression

Stephen J. Cox · 2002

A technique for adapting speaker-independent speech recognition models to the voice of a new speaker is presented. The technique is capable of estimating adapted parameters for all the speech models when only a small subset of the recognition vocabulary is spoken by the new speaker. Whereas previous methods have often assumed a transformation between the speaker-independent models and the adapted models, this technique models the relationship between different speech units using linear regression. The regression models are built off-line using the training-set data. At recognition-time, the speech models are adapted using the regression models and the new speaker's data, a procedure which is computationally cheap. Experimental results show a halving of the recognition error-rate when only about 8% of the vocabulary is given as enrollment data, and when half the vocabulary is given, a reduction in the error-rate of 78%.

Read the paper · More papers on PaperTik