Hybridization of Self Supervised Learning Models for Enhancing Automatic Arabic Speech Recognition
Hiba Adreese Younis, Yusra Faisal Mohammad · 2024
End-to-end systems for automatic recognition of Arabic speech have gained significant importance in the last few years. Arabic is considered as one of the low resource languages that lacks the labeled data required for training large models using supervised learning. To overcome this problem self-supervised learning models appears which can improve the robustness of the system by avoiding several issues with labels resulting from corrupted audio files. The objective of this paper is to assess the effectiveness of various models for Arabic speech recognition and also to compare the base SSL models with hybrid models in terms of word error rate(WER). The methodology includes hybridization of two state of the art models (wav2vec2xlsr, HUBERT) with adapter layer of MMS model. Another hybridization was also implemented by hybridization these SSL models with four layers of transformer encoder. The MMS model was also used for the first time for fine-tuning Arabic language using different hyper parameter values and gave competing results. As optimization algorithm, the SWATS mechanism was being used with these four hybrid models to improve their performance. Results showed that hybrid hubert_transformer model gave the best results on test dataset with 23.83 WER, followed by MMS model, hybrid wave transformer model, hybrid wave_MMS and hybrid hubert_MMS model respectively.