Enhancing Arabic Speech Emotion Recognition with Deep Learning
Mohamed Hassanain, Noran Shehata, Ahmed S.A. Mohamed, Maged Hammad, Sama Alqasaby, Ahmed Bayoumy Zaky · 2024
The development of Speech Emotion Recognition (SER) technology is crucial for improving artificial intelligence systems, enabling more natural and intuitive humancomputer interaction. Despite its potential, the advancement of SER systems in the Arabic language has been limited by the scarcity of comprehensive Arabic-language databases, constraining their effectiveness in Arabic-speaking markets. This study bridges this gap by enhancing Arabic SER technologies through the use of diverse datasets sourced from television shows, online talk shows, and other audio media. Four advanced deep learning models-parallel-CNN-attention-LSTM, parallel-CNN-transformer, Wav2Vec2-Large-xlsr-53-Arabic, and HuBERT-base-LS960-were employed, demonstrating significant advancements. Notably, the Parallel-CNN-LSTM model achieved an impressive accuracy of $93.81 \%$ on the BAVED dataset, outperforming other models, while the BAVED HuBERT model achieved an accuracy of $86 \%$. These results highlight the superiority of CNN-based architectures and pretrained models tailored for dataset characteristics. The findings advance the state-of-theart in Arabic SER, offering a robust foundation for developing more accurate systems tailored to Arabic language applications.