Pre-processing Analysis for Chinese Text Sentiment Analysis
Ang Li, Yunfang Chen · 2017
Nowadays, businesses and organizations are very eager to analyze emotional tendencies of the information generated by the consumers on the internet. Most of the current researches focus on the classification methods and vectorization of the text, there is a lack of research on the effects of pre-processing for Chinese text on sentiment analysis, especially for Chinese as a tonal language requires to deal with the problem of word segmentation. In this paper, we study the different effects of the eight short text pre-processing operations on the results of sentiment analysis, and we use the machine learning methods to carry on the sentiment classification. The classification methods include Naive Bayes, Support Vector Machine, Convolutional Neural Network and Long Short-Term Memory. The experimental results show that the effects of Chinese pre-processing and English pre-processing have some similarities. However, compared with english, the stop words in chinese contain more valuable information, and Convolutional Neural Network is more suitable for the case of fewer features.