A Fine Grained Sentiment Analysis of Arabic Language

Asim Tanveer, Mohibullah Khan, Rehan Sarwar, Naeem Aslam, Muhammad Fuzail · VAWKUM Transactions on Computer Sciences · 2024

This work focuses on fine-grained sentiment analysis of Arabic text using recent Natural Language Processing methods. Arabic is a language rich in variation, spoken by over 400 million people, yet there is a significant lack of resources for sentiment analysis. To address these challenges, this study employs AraBERT, a model specifically fine-tuned for Arabic text. A corpus of one hundred thousand Arabic reviews across categories such as hotels, books, and movies was scraped and cleaned. These reviews were then categorized into positive, negative, and mixed sentiments. AraBERT was compared with traditional machine learning methods, including Logistic Regression, Decision Tree, Naïve Bayes, and Random Forest. AraBERT achieved superior accuracy of 88\%, along with higher precision, recall, and F1 scores for both positive and negative sentiment classes compared to the other models. This work demonstrates that AraBERT effectively analyzes the syntactic and semantic structure of Arabic, making it a valuable tool for Arabic sentiment analysis across various applications. Future work will extend the model to handle neutral sentiments and include additional dialects to further improve its performance.

Read the paper · More papers on PaperTik