Combining a rule-based classifier with ensemble of feature sets and machine learning techniques for sentiment analysis on microblog
Umme Aymun Siddiqua, Tanveer Ahsan, Abu Nowshed Chy · 2016
Microblog, especially Twitter, have become an integral part of our daily life, where millions of users sharing their thoughts daily because of its short length characteristics and simple manner of expression. Monitoring and analyzing sentiments from such massive Twitter posts provide enormous opportunities for companies and other organizations to estimate the user acceptance of their products and services. But the ever-growing unstructured and informal user-generated posts in Twitter demands sentiment analysis tools that can automatically infer sentiments from Twitter posts. In this paper, we propose an approach for sentiment analysis on Twitter, where we combine a rule-based classifier with a majority voting based ensemble of supervised classifiers. We introduce a set of rules for the rule-based classifier based on the occurrences of emoticons and sentiment-bearing words. To train the supervised classifiers, we extract a set of features grouped into Twitter specific features, textual features, parts-of-speech (POS) features, lexicon based features, and bag-of-words (BoW) feature. A supervised feature selection method based on the chi-square statistics (χ2) and information gain (IG) is applied to select the best feature combination. We conducted our experiments on Stanford sentiment140 dataset. Experimental results demonstrate the effectiveness of our method over the baseline and known related work.