Confidence-driven iterative speaker adaptation in transcription-mode speech recognition
Jiazhi Ou, Kaijiang Chen, Zongge Li · 2002
Speaker adaptation in transcription-mode speech recognition means adaptation on the test set itself rather than on an independent data set. A confidence-driven iterative adaptation approach is presented. Speech recognition and speaker adaptation are performed alternately. A combination of MLLR (maximum likelihood linear regression) and MAP (maximum a posteriori) approaches is used for adaptation. To avoid erroneous transcription in adaptation, an N-best based confidence measure is introduced. The procedure includes a coarse classification and a delicate measurement. Experiments were carried out on a large vocabulary Mandarin continuous speech recognition. A baseline model for comparison was built. The experimental results showed that our proposed approach reduced the word error rate by 51.8% compared to the initial speaker-independent system and 39.3% compared to the baseline model.