Enhancing Arabic ASR in Noisy and Transcoding EVS Conditions: A Multimodal Deep Learning Study

Lallouani Bouchakour, Khaled Lounnas, Ahmed Krobba · Annals of Computer Science and Information Systems · 2025

In this paper, we investigate the impact of speech transcoding and noise on the performance of Arabic automatic speech recognition (ASR) systems based on deep learning.We apply Non-negative Matrix Factorization (NMF) as a denoising preprocessing step to enhance robustness to noise.Three deep architectures-CNN-LSTM, LSTM, and DNN-are evaluated using fused acoustic features including MFCCs, Mel-spectrograms, and Gabor filter representations.Experiments are conducted under four signal-to-noise ratio (SNR) conditions (-5 dB, 0 dB, 5 dB, and 10 dB) on both transcoded and nontranscoded speech.Results show that the CNN-LSTM model achieves the highest accuracy of 87% at 10 dB SNR on clean (non-transcoded) speech using multimodal features.However, speech recognition performance degrades by 2-4% when using the Enhanced Voice Services (EVS) codec, especially in highnoise environments.Specifically, accuracy drops from 65.00% to 61.43% at -5 dB SNR, and from 87.00% to 84.00% at 10 dB SNR due to transcoding.These findings highlight the negative impact of mobile codec compression on ASR systems, particularly under low-SNR conditions.Our study confirms the effectiveness and stability of NMF-based feature fusion and denoising in improving recognition, offering insights into deploying Arabic ASR in real-world scenarios such as mobile and VoIP communications.

Read the paper · More papers on PaperTik