Sentiment classification of short texts based on semantic clustering

Yunchao He, Chin‐Sheng Yang, Liang-Chih Yu, Kuo-Hua Robert Lai, Weiyi Liu · 2015

Due to the popularity and ubiquity of social networks, sentiment analysis has become an important and well-covered research area. Short texts usually encounter sparsity problems in representations for their limited texts length. We address this issue by clustering short texts to form a long text. Specifically, we first use k-means clustering algorithms to form k clusters, where each cluster contains texts having close semantic similarity with the same sentiment polarity. Then, these clusters are used to train classifier. In evaluation, the unlabeled text is merged to the most similar positive and negative clusters respectively, and its sentiment polarity is determined by the change of the two clusters' probabilistic estimates. Experiments on the Twitter dataset show that our approach can achieve a significantly better performance than the bag-of-words method.

Read the paper · More papers on PaperTik