Enhancing Moroccan Dialect Named Entity Recognition through Data Augmentation with AraGPT-2 and Similarity Measures

Manel Affi, Chiraz Latiri · Procedia Computer Science · 2024

The effectiveness of learning models is significantly influenced by the accessibility and sufficiency of training data. To address the challenge of dataset adequacy, researchers have extensively explored data augmentation (DA) as a promising strategy. DA involves generating new data instances through applied transformations to existing data, thereby augmenting the dataset in terms of both size and variability. This method has demonstrated success in enhancing model performance and accuracy across a spectrum of natural language processing (NLP) tasks. Despite its proven efficacy, there has been limited exploration of DA specifically for the Arabic language, with existing studies relying predominantly on conventional methods such as paraphrasing or techniques based on introducing noise. In this study, we introduce an innovative approach to Arabic DA, tailored for the Moroccan dialect. We leverage the advanced modeling technique AraGPT-2 for the augmentation operation. The generated sentences are assessed based on context, semantics, diversity, and novelty, utilizing metrics such as Euclidean distance, cosine similarity, Jaccard index, and BLEU score. Subsequently, we employ the Multi-dialect-Arabic-BERT transformer for the Named Entity Recognition (NER) task, evaluating its performance on the augmented Moroccan dialect dataset. The experiments were conducted on the DarNER-corpus, representing the Moroccan dialect. The outcomes illustrate that our proposed methodology significantly enhances the NER task for the Moroccan dialect, showcasing a noteworthy increase in F1-score by 2.89 points.

Read the paper · More papers on PaperTik