Research on End-to-End Pronunciation Error Detection Based on Transfer Learning
文明 高 · Computer Science and Application · 2021
自动发音检错是为了满足第二语言学习者发音练习的需求,而先进的自动发音检错系统通常取决于声学模型识别率。随着深度学习技术的发展,端到端声学模型算法已经逐渐成熟,为发音检错算法研究提供的新思路。本文首先构建了基于连接时序分类(Connectionist Temporal Classification, CTC)算法的端到端发音检错声学模型架构。其次,基于二语迁移现象,L2发音往往带有其母语的音素特征,本文利用迁移学习算法提高基于母语的声学模型性能,从而提高发音检错准确率。通过迁移中文母语音素特征的声学模型相比于只使用英文母语的声学模型在错误音素率上有所下降,并且训练时间减少了7.3%。在发音检错性能上检错正确率提升了2.06%。 Automatic pronunciation error detection is to meet the needs of second language learners’ pronunciation practice, and advanced automatic pronunciation error detection systems usually depend on the recognition rate of the acoustic model. With the development of deep learning technology, end-to-end acoustic model algorithms have gradually matured, providing new ideas for the research of pronunciation error detection algorithms. This paper first builds an end-to-end pronunciation error detection acoustic model architecture based on the Connectionist Temporal Classification (CTC) algorithm. Secondly, based on the phenomenon of second language transfer, L2 pronun-ciation often has the phoneme characteristics of its native language. This paper uses transfer learning algorithms to improve the performance of the acoustic model based on the native language, thereby improving the accuracy of pronunciation error detection. Compared with the acoustic model that only uses the native English language, the acoustic model that transfers the Chinese phoneme features has a lower error phoneme rate, and the training time is reduced by 7.3%. The correct rate of error detection in pronunciation error detection performance has increased by 2.06%.