A Comparative Evaluation of Word Embeddings Techniques for Twitter Sentiment Analysis
Ibrahim Kaibi, El Habib Nfaoui, Hassan Satori · 2019
The writer's opinion is important information, which is becoming increasingly desirable with the increasing volume of user-generated content on the Web. That's why made sentiment analysis an important tool for extracting this information, which can usually be positive, negative or neutral. To target sentiment classification problem, the standard approach used is the binary classification, considering the sentiment (or the polarity), positive or negative. The classification result depends on the text representation and then the extracted features used to train the classifier. Word embeddings techniques have emerged as a prospect for generating word representation for different text mining tasks, especially sentiment analysis. In this paper, we focus on the comparison of three commonly used word embeddings techniques (Word2vec, Fasttext and Glove) on Twitter datasets for Sentiment Analysis, employing six popular machine learning algorithms, namely, GaussianNB, LinearSVC, NuSVC, LogisticRegression, SGD and RandomForest. We find that Fasttext representation used with NuSVC, which is a type of SVM classifier, outperforms the other combinations in accuracy.