Phoneme-Level Pronunciation Error Detection and Classification Using Deep Belief Networks and Support Vector Machines
Shaonan Lin, Yanshun Feng · Journal of Circuits Systems and Computers · 2025
With the rise of online learning, computer-assisted language learning (CALL) has become an increasingly popular choice for language learners. Among the core components of CALL, pronunciation error detection and diagnostic feedback play a crucial role in analyzing learners’ pronunciation issues and providing corrective suggestions, ultimately improving pronunciation proficiency and learning efficiency. While most existing studies focus solely on error detection, feedback correction remains underexplored. This paper addresses phoneme-level pronunciation errors caused by nonstandard articulatory movements, specifically focusing on six error types: Rising, Lowing, Fronting, Backing, Lengthing and Shorting. Using a machine learning approach, we develop a classification error detection model based on deep belief networks (DBN) and support vector machines (SVM), integrating the One-Class SVM (OC-SVM) approach to tackle issues of imbalanced sample data. The proposed DBN-SVM model is capable of detecting three additional error types — Fronting, Backing and Lengthing — completing the detection of all six error types. Experimental results validate the effectiveness of the model in pronunciation error detection and classification.