Arabic speech recognition based on a CNN-BLSTM combination

Rafik Amari, Abdelkarim Mars, Mounir Zrigui · 2022

Despite advances in speech recognition technology, Arabic speech recognition remains largely unsolved due to its many difficulties and challenges. The performance of the best existing recognizers is much lower than those developed in English. Deep Neural Networks (DNNs) have shown excellent performance in acoustic modeling for speech recognition. In this work, a new discontinuous Arabic speech recognition model is proposed. It associates a deep convolutional neural network (CNN) architecture with a long-term bi-directional memory (BLSTM). The optimal network structure and training strategy for the model are examined. The Arabic Speech Corpus of Isolated Words (ASDS) and the Spoken Arabic Digits (SAD) database were used for all experiments. The results demonstrate the strength and benefits of the CNN-BLSTM method, which provides the best detection accuracy.

Read the paper · More papers on PaperTik