Transliteration-Aided Transfer Learning for Low-Resource ASR: A Case Study on Khalkha Mongolian
Dalai Mengke, Yan Meng, Péter Mihajlik · Electronics · 2025
Automatic Speech Recognition (ASR) systems have made consistent advancements, achieving notable improvements in state-of-the-art performance across various languages. However, their effectiveness often declines significantly in low-resource settings, where data and linguistic resources are limited. This paper addresses the challenges of ASR for a low-resource language, Khalkha Mongolian, by leveraging a transliteration-aided transfer learning approach. Specifically, it improves the ASR system for Khalkha Mongolian by transliterating text from a well-resourced Chakhar Mongolian (Uighur script) dataset to the Cyrillic script and then fine-tuning it with Khalkha Mongolian data. The method effectively enhances the ASR performance of Khalkha Mongolian. The effectiveness of the proposed method was validated on three popular ASR models, Wav2Vec2-BERT, Conformer-Large, and Whisper-large-v3. Among these models, the best relative improvement in word error rate (WER) reaches 32.50%, while the absolute improvement reaches 19.26%.