Research on Pre-trained Speech Recognition Models for Pronunciation Improvement
Kanev I. Anton, Alexey V. Papin · 2025
The main objective of this work is to study the effectiveness of various speech recognition models in Russian using the metrics Word Error Rate (WER), Message Error Rate (MER), Word Information Loss (WIL), and Character Error Rate (CER). The work is dedicated to researching the application of machine learning technologies, particularly pre-trained models from Hugging Face, for analyzing and correcting speech defects. The relevance of the topic is due to the growing need for speech defect correction and diction improvement, especially for students at the Bauman Moscow State Technical University, which hosts a center for students with disabilities.