Mizo Automatic Speech Recognition: Leveraging Wav2vec 2.0 and XLS-R for Enhanced Accuracy in Low-Resource Language Processing

Andrew Bawitlung, Sandeep Kumar Dash, Radha Mohan Pattanayak · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

This study introduces a Mizo Automatic Speech Recognition (ASR) approach by fine-tuning Wav2vec 2.0 and XLS-R models. The research presents the newly developed Mizo speech dataset, MiZonal v1.0 which significantly contributes to the advancement of low-resource language processing and plays a crucial role in preserving the Mizo language, thereby enhancing the training and assessment of speech models for this underrepresented language. It focuses on evaluating the effectiveness of these models in handling Mizo speech data, with particular emphasis on their performance in converting numerical numbers into Mizo cardinal words, which have a positive effect on the Word Error Rate (WER). The findings reveal that while the Wav2vec-Base-Mizo-Lus model achieved a WER of 16.59%, the XLS-R-300M-Mizo-Lus model outperformed it significantly, achieving a WER of 11.84% and setting a new benchmark for accuracy in the Mizo language. This shows the importance of using large multilingual speech recognition models and the cross-lingual abilities of models such as XLS-R are essential for ASR tasks in low-resource languages, leading to progress in Mizo speech technology and its applications.

Read the paper · More papers on PaperTik