Improving Low Resources Arabic Speech Recognition using Data Augmentation

Mohanad Khudhair, Ahmed Talib · 2022

The use of data-augmentation on training data can significantly improve the robustness of the deep neural networks-based automatic speech recognition (ASR) system. We propose an approach to building and developing an ASR system for the low-resource Arabic language using an End-to-End model and demonstrates the impact of data augmentation and the suggested language model on the results. We aimed to develop a system that transcribes Arabic audio containing human speech into text since a few existing systems tailored for Arabic are developed compared to other languages. Our model is trained based on the DeepSpeech2 framework that uses End-to-End deep learning. When data augmentation is utilized during training, the word error rate (WER) is improved by around 13%, and when using the language model, the word error rate is reduced by around 48%. Our best model achieves a competitive word error rate when the system was evaluated on the Common Voice 8.0 dataset at 2.8 of WER.

Read the paper · More papers on PaperTik