Customer Sentiment Analysis in Arabic Social Media

Ayman Yafoz · 2020

Performing sentiment analysis on text document is an active research area. Social media includes valuable information resources in various languages, which encompass reviews, comments, tweets, posts, opinions, articles and other text resources. These could be analysed to explore people’s opinions, attitudes, emotions and sentiments toward various subjects and commodities. Hence, this thesis targets customer sentiment analysis in Arabic social media, with a focus on both the real estate and automobile industries. In this regard, automated analyzer systems were proposed in this thesis, to classify the sentiment polarity of each social media customer feedback review into “positive”, “negative” or “mixed”. This was achieved by gathering data from online customers who wrote reviews about real estate or automobiles. These online forums are considered among the largest customer social media discussing feedback on real estate and automobiles in Gulf Cooperation Council (GCC) dialects, in particular, and in the Arabic language in general. The datasets are in both GCC dialects and Modern Standard Arabic (MSA). Moreover, they were annotated using three annotators, and inter-rater agreement was calculated to assess both the consistency and reliability of the annotators. Following this process, the proposed systems performed a series of preprocessing operations on the collected data, with the purpose of both cleaning and preparing them for classification. The normalization process included tokenization, performance of regular expression processes, lemmatization, and sentence segmentation. Moreover, it also encompassed a cleansing process in order to have a noise-free dataset. Furthermore, feature selection was performed. Part-of-speech tagger was adopted to enhance the classification process through improving sentiment words recognition by the classifier. TF-IDF (Term Frequency-Inverse Document Frequency) was another feature selection procedure used to reflect the term’s importance and informativeness in a particular document. Following that, the n-grams feature was utilized to generate features for the classifiers. Afterwards, the datasets were divided into training and testing datasets, whereas the cross-validation method was applied to randomly split the training dataset in a non-overfitting manner. Then the case of imbalanced datasets was handled so as to gain a sufficient ensemble of minority categories and enhance the performance of the classifier. Moreover, a set of machine and deep learning classifiers were used to categorize data, and the hyperparameters were tuned in order to achieve higher classification performance and results. Furthermore, a set of information visualization techniques was adopted to view the results, depict the performance of the classifiers as well as show the most important terms that had affected the classification process. Finally, the results were shown and future suggestions were laid out to enhance these results, and the performance of the proposed systems were provided. Despite suggesting future improvements, the results are competitive compared to those achieved through other contributions addressing Arabic sentiment analysis.

Read the paper · More papers on PaperTik