Sentiment Analysis Using Learning Approaches Over Emojis for Turkish Tweets
Rıza Velioğlu, Tuğba Yıldız, Savaş Yıldırım · 2018 3rd International Conference on Computer Science and Engineering (UBMK) · 2018
With the rise of the usage and interest on social media platforms, emojis have become an increasingly important part of the written language and one of the most important signals for micro-blog sentiment analysis. In this paper, we employed and evaluated classification models using two different representations based on bag-of-words and fastText to address the problem of sentiment analysis over emojis/emoticons for Turkish positive, negative and neutral tweets. At first, the bagof-words approach is used as a simple and efficient baseline method for tweet representation, where the classifiers such as Naive Bayes, Logistic Regression, Support Vector Machines, Decision Trees have been applied to these tweets. Secondly, we utilized fastText to represent tweets as word n-grams for sentiment analysis problem. The results show that there is no significant difference between the two models. While fastText shows 79% and the Linear Regression classifier obtains 77% F1-score for binary classification, fastText performs 62% and Linear Regression has 58% F1-score for multi-class classification. This study is considered as the first study that contributes to the literature by applying different vector representations such as bag-of-words and fastText to predict Turkish tweets over emojis. This study can also be utilized to predict emojis on social media context in the future.