AraABSAMD: A novel arabic dataset for aspect-based sentiment analysis in the Moroccan education domain

Lachhab Youssef, Ziyati Elhoussaine · Array · 2026

The Arabic language suffers from a lack of annotated datasets in aspect-based sentiment analysis compared to other languages, such as English. It requires high-quality annotated datasets to enhance model accuracy, particularly in niche areas such as education. This paper enriches the Arabic dataset by establishing the first annotated dataset in the field of education, AraABSAMD. It includes 1360 Twitter reviews from the Moroccan community, written exclusively in modern standard Arabic (MSA), covering five aspect categories: Curriculum and Instruction, Student Experience, Faculty and Staff, Outcomes and Performance, and Technology Integration, with the four sentiment polarities — positive, negative, neutral, and conflict. This paper presents an annotation platform for aspect-based sentiment analysis and describes the various techniques and methodologies employed for the practical annotation of the dataset. It details the features of the developed platform, which include user-friendly annotation tools, and dataset structuring. Moreover, it discusses the techniques employed for annotation, including manual labeling and reporting inter-annotator agreement scores to ensure reliability. Furthermore, it describes challenges with the annotation process, such as handling ambiguous sentiments and ensuring consistency across different annotators, and outlines strategies to address these challenges. The dataset was evaluated using three models: bert-base-arabertv2, ARBERT, and camelbert-msa. For the aspect category detection (ACD) task, bert-base-arabertv2 and camelbert-msa achieved the highest performance, each attaining an F1-score of 69%. In the aspect category polarity (ACP) task, ARBERT performed best with an F1-score of 76%. Finally, for the combined aspect term extraction and aspect sentiment classification (ATE + ASC) task, ARBERT achieved an F1-score of 62%.

Read the paper · More papers on PaperTik