A Guide to Cost-Effective Collection of Colloquial Algerian Arabic Speech Data

Touahmi Moussa, Ouamane Abdelmalik, Chouchane Ammar · 2024

Modern Standard Arabic (MSA) serves as the official language across all Arab countries, employed in administrative, educational, official broadcast, and press settings. However, in everyday informal communication, each Arab country adopts its own variant dialect, commonly known as colloquial dialect. This colloquial form has emerged as the dominant language on social media platforms due to the widespread usage of such platforms over the past decade. Consequently, there is a growing demand for language resources and natural language processing (NLP) systems that cater to these dialects. This paper details the development of an Algerian Arabic dialect dataset, with a focus on average colloquial dialect to ensure effective performance on standard Algerian dialect. Carefully curated, the dataset encompasses a diverse collection of high-quality audio extracted from social media platforms. Employing a finetuning approach, we leverage existing Automatic Speech Recognition (ASR) systems, particularly the Whisper model by OpenAI for its noise robustness and adaptability to varying audio quality. The paper delves into the process of efficiently collecting speech data, while also outlining potential applications and future directions. Researchers in speech analysis and natural language processing will find this work invaluable for advancing their knowledge of Algerian Arabic and fostering accurate ASR system development tailored to the dialect.

Read the paper · More papers on PaperTik