Arabic Speech Recognition based on Self Supervised Learning

Hiba Adreese Younis, Yusra Faisal Mohammad · 2023

Automatic Arabic Speech Recognition (AASR) has gained significant attention in recent years due to its potential applications in various fields such as transcription, voice assistants, and language learning. However, ASR systems typically require a large amount of labelled data for training, which can be expensive and time-consuming to obtain, especially for languages with limited resources like Arabic. To address this challenge, self-supervised learning techniques have emerged as a promising approach to improve ASR performance by leveraging unlabeled data.In this research, the architecture of the HUBERT model was enhanced by removing the last layer and adding additional layers including (normalization layer, two linear layer followed by RELU activation function, drop out layer), and finally a linear projection layer for alignment with the vocabulary size of Arabic language dataset which led to better performance and smaller word error rate(WER) as compared to HUBERT base model. Many experiments were also applied in this research for enhancing mono and multi lingual ASR models using different combination of hyper parameters through trial and error. Results showed that multi lingual wav2vec2XLSR model outperformed monolingual HUBERT model and massive multilingual speech(MMS) model fine-tuned on Arabic common voice dataset. also, the proposed modified Hubert model with specific type of learning rate scheduler(MHLR) gave better results than base HUBERT model and good results as compared with large MMS model. These models were evaluated using WER and WRR metrics and showed 40% for wav2vec2xlsr model and 42%, 45%, 54%, for MMS, MHLR and HUBERT base models respectively.

Read the paper · More papers on PaperTik