Implementation of Machine Learning Algorithms in Arabic Sentiment Analysis Using N-Gram Features

Donia Gamal, Marco Alfonse, El-Sayed M. El-Horbaty, Abdel-Badeeh Mohamed Salem · Procedia Computer Science · 2019

Sentiment analysis (SA) is a scholarly process of extricating and classifying individuals’ emotions and feedbacks expressed in source text content. It is one of the pursued subfields of Computational Linguistics (CL) and Natural Language Processing (NLP). The evolution of social media based applications has generated a big amount of personalized reviews of different related information on the Web in the form of tweets, status updates, and many others. Several approaches have come into the spotlight in recent years to accomplish SA, the most part of SA researches have been applied utilizing the English language. SA in Arabic online social media may be slacking behind commonly because of the difficulties with handling the morphologically complex Arabic natural language and the lack and absence of accessible tools and assets for extracting Arabic opinions from the text. This research is aimed to analyze the collected twitter posts in different Arabic Dialects and a comparison between the various algorithms used for SA with various n-gram as a feature extraction method. The measurement of the performance of different algorithms is evaluated in terms of recall, precision, f-measure, and accuracy. The experiment results show that unigram with Passive Aggressive (PA) or Ridge Regression (RR) gives the highest accuracy 99.96 %.

Read the paper · More papers on PaperTik