Indonesian speech recognition based on Deep Neural Network

Ruolin Yang, Jian Yang, Yu Lu · 2021

The speech recognition system using deep neural network and large-scale datasets have had good performance. Indonesian is a language with hundreds of millions of people. Due to the late start of research and the lack of large enough speech data, the research of Indonesian speech recognition lags. In view of this situation, the deep neural network model is used to study Indonesian speech recognition. According to the transferability of deep neural network and the similarity between English and Indonesian, three transfer learning methods to improve the performance of Indonesian speech recognition system are proposed, and the effectiveness of different transfer methods is verified by experiments, and the acoustic model training scheme is further optimized. The experimental results show that the TDNN-Attention-HMM model based on hierarchical weight transfer has better performance than the other two transfer models. In this paper, this model is used as the optimal acoustic model of Indonesian speech recognition, and the word error rate is 6.79%. Compared with the DNN-HMM baseline system, the word error rate is relatively reduced by 26.52 %.

Read the paper · More papers on PaperTik