A comparison of features extraction methods for Arabic sentiment analysis

Mohammed Kasri, Marouane Birjali, Abderrahim Beni‐Hssane · 2019

Natural Language Processing (NLP) has built up so much importance in the past few years. With machine learning, NLP can detect a lot of unseen information from a huge volume of textual data, which can be helpful for sentiment analysis and text classification. For those tasks, features can be extracted, the operation bases on extracting an important subset of features from a data. However, identifying convenient features is very important for improving the NLP tasks. For Arabic Language the process of getting the related features is difficult due to many reasons, for example, this language has the most words compared to other languages. Our contribution in this paper is to analyze the impact of feature extraction methods such as Bag-of-Words, TF-IDF, and word2vec on the performance of sentiment analysis using Arabic language. The extracted features are evaluated using many machine learning algorithms like Logistic Regression and Support Vector Machine.

Read the paper · More papers on PaperTik