Analysis of Moroccan Arabic Speech Recognition for Integration in Voice Command Systems
Abderrahim Ezzine, Naouar Laaidi, Hassan Satori · 2024
Automatic speech recognition (ASR) systems transform acoustic waveforms into corresponding text representations, effectively capturing the semantic content conveyed by speakers. This paper presents the development of an ASR system tailored to Moroccan Arabic (Darija) for recognizing connected words in command-and-control scenarios. The system was trained on a dataset consisting of 2,394 words, each repeated 10 times, articulated by 19 speakers (both male and female) in Moroccan Arabic. It leverages the CMU Sphinx toolkit and employs Hidden Markov Models (HMMs) integrated with Gaussian Mixture Models (GMMs) for acoustic modeling. Mel-Frequency Cepstral Coefficients (MFCCs) were used for feature extraction. Experimental results demonstrate a recognition accuracy of $\mathbf{9 3 . 3 3 \%}$ for connected words, achieved using a configuration with four Gaussian mixtures.