Iterative Deep Learning-based Acoustic Models using Transcription Agreement from Multi-Models Automatic Speech Recognitions

Iftitakhul Zakiah, Dessi Puji Lestari · 2020

Compared to English or languages with a lot of resources, Indonesian is one of the languages which is categorized as under resource language. Indonesian has a limited transcribed speech corpus that can be used to build an acoustic model of a speech recognition system with a supervised learning approach. To overcome this problem, this research proposes the development of iterative acoustic models using an additional unlabeled speech corpus. We used unlabeled data to rebuild the acoustic models by using segment's transcriptions generated by the previously supervised developed ASR. To get more reliable transcription, we used four ASRs with four types of deep learning-based acoustic models (DNN, LSTM, CNN, and TDNN) and selected segments with consistent transcripts provided by all models or segments with fully agreement labels. This technique relatively improves the accuracy of the three ASRs that were developed using DNN-based acoustic models by 1.95%, CNN by 1.56%, and TDNN by 2.59%.

Read the paper · More papers on PaperTik