Continuous Arabic Speech Recognition Model with N-gram Generation Using Deep Speech

Fawaz S. Al–Anzi, S.T. Bibin Shalini · 2024

A speech recognition system aims to translate audio input into a string of words. Deep Speech is an end-to-end programmed communication acknowledgment engine that has shown promising effects in speech recognition. In this work, Mozilla's Deep Speech framework shows effectiveness for Arabic speech recognition. Baidus DeepSpeech's neural network architecture comprises several layers of recurrent neural networks (RNNs), and RNN is its fundamental component. It is a speech-to-text technology that transcribes spoken language into written text by utilizing deep learning algorithms. To understand the complex patterns and interactions between spoken sounds and their corresponding textual representations, these RNNs are trained on a large dataset of voice and alphabet text. Gathered transcripts and audio datasets in Arabic for testing or training. Generated acoustic and language models for the better training of Arabic audio data sets. To complete this task, a virtual environment is set up and launched, and a docker file is prepared using PyCharm in Python 3.6 version. Auditory records were transformed to a comparable WAV design. Spectrograms and MFCCs were used for feature extraction and preprocessing. Appropriate indicators, including word mistake rate, character error rate, and loss, are used to assess the model's performance. The Deep Speech is used to achieve high-quality Arabic speech recognition with reduced loss, word error, and Character error rate. Baidus deep speech aims to perform well in both isolated and end-to-end speech recognition. It finds the global optimum value with good precision and a low error rate in a reasonable period. The accuracy of Arabic automatic voice recognition is increased using this approach. This work can lead to various future directions for improving accuracy in Arabic speech recognition.

Read the paper · More papers on PaperTik