Word-selective training for speech recognition

Terri Kamm, Gérard G. L. Meyer · 2004

We previously proposed (Kamm and Meyer (2001, 2002)) a two-pronged approach to improve system performance by selective use of training data. We demonstrated a sentence-selective algorithm that, first, made effective use of the available humanly transcribed training data and, second, focused future human transcription effort on data that was more likely to improve system performance. We now extend that algorithm to focus on word selection, and demonstrate that we can reduce the error rate from 10.3 % to 9.3 % on a simple, 36-word corpus, by selecting 30 % (15 hours) of the 50 hours of training data available in this corpus, without knowledge of the true transcription. We also discuss application of our word selection algorithm to the Wall Street Journal 5 K word task. Preliminary results show that we can select up to 60 % (48 hours) of the training data, with minimal knowledge of the true transcription, and match or beat the error rate of a system built using the same amount of randomly selected training data.

Read the paper · More papers on PaperTik