A machine Learning approach for sentiment analysis in the standard or dialectal Arabic Facebook comments
Abdeljalil Elouardighi, Mohcine Maghfour, Hafdalla Hammia, Fatima-zahra Aazi · 2017
Social networks like Facebook contain an enormous amount of data, called Big Data. Extracting valuable information and trends from these data allows a better understanding and decision-making. In general, there are two categories of approaches to address this problem: Machine Learning approaches and lexicon based approaches. This work deals with the sentiment analysis for Facebook's comments written and shared in Arabic language (Modern Standard or Dialectal) from a Machine Learning perspective. The process starts by collecting and preparing the Arabic Facebook comments. Then, several combinations of extraction (n-grams) and weighting schemes (TF / TF-IDF) for features construction are conducted to ensure the highest performance of the developed classification models. In addition, to reduce the dimensionality and improve the classification performance, a features selection method is applied. Three supervised classification algorithms have been used: Naive Bayes, Random Forests and Support Vectors Machines using R software. Our Machine Learning approach using sentiment analysis was implemented with the purpose of analyzing the Facebook comments, written in Modern Standard Arabic or in Moroccan Dialectal Arabic, on the Morocco's Legislative Elections of 2016. The results obtained are promising and encourage us to continue working on this subject.