Multi-Domain Dataset for Moroccan Arabic Dialect Sentiment Analysis in Social Networks

Sara El Ouahabi, Safâa El Ouahabi, El Wardani Dadi · 2023

Social media generates a massive amount of textual data every day, users post comments, tweets, messages, reviews, and ratings, creating very large and diverse data. The comments are often very authentic, as users express their opinions freely and without influence. This can help capture people's true feelings about a topic in real-time. Currently, the majority of research and projects primarily focus on Indo-European languages, particularly English, leaving a significant portion of dialect-speaking communities underserved and overlooked. In this work, we focus our study on Moroccan dialect sentiment analysis in different social networks; we have collected and manually annotated a rich dataset from different social media (Facebook, Twitter, YouTube, Instagram, and websites. To the best of our knowledge, there is currently no publicly accessible sentiment analysis dataset that encompasses all social networks and specifically caters to the sentiment analysis task. Furthermore, the dataset we have collected stands out as the largest dataset focused on sentiment analysis specifically for Morocco. It is characterized by its size (70 332), quality, variety (covering a lot of domains), novelty (it contains all topics and trends from 2020 to now) Keywords─ Moroccan Arabic dialect, sentiment analysis, dataset, social media.

Read the paper · More papers on PaperTik