Unlocking ASR Potential for Under-Resourced Languages: A Comparative Study of Semi-Supervised Approaches for Accurate Unlabeled Data Transcription

Abdullah Mahmoud, Ahmed S. ELSayed, Zaki Taha Fayed · 2025

Automatic Speech Recognition (ASR) enhances human-computer interaction and accessibility across industries, but developing robust systems for under-resourced languages like Arabic and specialized domains is challenging due to the difficulty of obtaining high-quality manual transcriptions. This paper compares two semi-supervised transcription techniques: confidence-based and model-ensembling methods for generating transcriptions for Arabic speech data. The confidence-based method generates transcriptions with a baseline ASR model, calculates confidence scores, and selects samples above a threshold. The model-ensembling approach uses two complementary ASR models, hybrid and end-to-end, aligning outputs via Levenshtein distance to extract accurate transcriptions. Using the Arabic Multi-Genre Broadcast (MGB2) dataset, it is found that the model-ensembling technique produces higher-quality labeled data than the confidence-based method, as reflected in the Weighted Average Relative WER Improvement scores (12.56 vs. 8.71). Notably, ASR models trained on automatically labeled data performed similarly to those trained on manually labeled data, achieving 12.56 and 12.37, respectively. This study demonstrates the potential of automatic transcription methods to enhance ASR model training and accessibility.

Read the paper · More papers on PaperTik