Transfer Learning for Less-Resourced Semitic Languages Speech Recognition: the Case of Amharic

Yonas Demeke Woldemariam · DiVA at Umeå University (Umeå University) · 2020

While building automatic speech recognition (ASR) requires a large amount of speech and text data, the problem gets worse forless-resourced languages. In this paper, we investigate a model adaptation method, namely transfer learning for a less-resourced Semiticlanguage i.e., Amharic, to solve resource scarcity problems in speech recognition development and improve the Amharic ASR model. Inour experiments, we transfer acoustic models trained on two different source languages (English and Mandarin) to Amharic using verylimited resources. The experimental results show that a significant WER (Word Error Rate) reduction has been achieved by transferringthe hidden layers of the trained source languages neural networks. In the best case scenario, the Amharic ASR model adapted fromEnglish yields the best WER reduction from 38.72% to 24.50% (an improvement of 14.22% absolute). Adapting the Mandarin modelimproves the baseline Amharic model with a WER reduction of 10.25% (absolute). Our analysis also reveals that, the speech recognitionperformance of the adapted acoustic model is highly influenced by the relatedness (in a relative sense) between the source and thetarget languages than other considered factors (e.g. the quality of source models). Furthermore, other Semitic as well as Afro-Asiaticlanguages could benefit from the methodology presented in this study.

Read the paper · More papers on PaperTik