Speech recognition for low-resource languages using large language models and related-language data

Jiqiao Zhang, Degen Huang · 2025

With the widespread adoption of large language models (LLMs) in natural language processing, integrating them with automatic speech recognition (ASR) has become a mainstream approach. However, existing research primarily focuses on high-resource languages, with limited effectiveness for low-resource languages due to model constraints and data scarcity. To address this, we propose a three-layer model for low-resource ASR, consisting of a speech encoder, an intermediate adapter, and an LLM, enabling LLM adaptation to low-resource languages. To mitigate data limitations, we employ a two-stage training strategy to enhance speech-text alignment and leverage similar-language data. Experiments show that our model reduces error rates by 0.5%–1.5% on Chinese, Uyghur, Hindi, and Thai compared to fine-tuning pretrained models, with further improvements when trained on same-language-family data, demonstrating its effectiveness in low-resource ASR.

Read the paper · More papers on PaperTik