Using untranscribed training data to improve performance

George Zavaliagkos, Man-Hung Siu, Thomas Colthurst, Jayadev Billa · 1998

This paper explores techniques for utilizing untranscribed training data pools to increase the available training data for automatic speech recognition systems. It has been well estab-lished that current speech recognition technology, especially in Large Vocabulary Conversational Speech Recognition (LVCSR), is largely language independent, and that the dominant factor with regards to performance on a certain language is the amount of available training data ([4]). The paper addresses this need for increased training data by presenting ways to use untranscribed acoustic data to increase the training data size and thus improve speech recognition. 1.

Read the paper · More papers on PaperTik