Text Corpus Augmentation to Represent Filled Pause in Indonesian Spontaneous Speech Recognition System
Candra Bella Vista, Dessi Puji Lestari, Dwi Hendratmo Widyantoro · 2019
Filled pause is one of the characteristics of spontaneous speech which causes a decrease in recognition accuracy. Including filled pause into language model is one of possible solution. In this case, the limitation of the Indonesian spontaneous text corpus is the main problem. In this paper, we employed augmentation on the non-spontaneous written text by inserting filled pause into non-spontaneous text corpus. Augmentation to the text corpus that will be used to build language model, has an impact on perplexity and WER reduction, about 11,1 % and 2,3% respectively. The augmentation strategy by adding filled pause based on hidden event language model estimation gives the best results both in the perplexity of language model and performance of speech recognition.