Arabic Dialect Speech-Text Recognition Using Deep Learning
Omar Alboom, Manar Alkhatib, Ashwaq Faisal, Khaled F. Shaalan · 2025
This study explores the challenges of Arabic speech recognition, focusing on deep learning-based transcription models due to the language’s complexity, limited audio datasets, and regional variations. Traditional models struggle with accurate transcription, prompting an evaluation of TestRCNN and Hybrid TestRCNN-CNN Models. The research involves preprocessing Arabic speech data, extracting features using MFCCs, and encoding labels for model training. The TestRCNN Model, integrating convolutional and recurrent layers, achieves 93% accuracy with a WER of 0.0986 but struggles with closely related phonetic sounds. To enhance performance, a hybrid approach combining CNN’s spatial feature extraction with TestRCNN’s temporal learning is introduced. The Hybrid TestRCNN-CNN Model outperforms its counterpart, achieving 94% accuracy and significantly reducing WER to 0.0460. The study provides a comparative analysis of these models, demonstrating the hybrid model’s effectiveness in Arabic speech recognition, and offering valuable insights for future advancements.