Multi-classification of microblog’s comments based on feature combination
Li Bing Bai · 2021
Text classification algorithm develops rapidly in recent years, especially with the promotion of deep learning, the accuracy of various algorithms has reached a certain height. However, an obvious drawback of deep learning is that a large number of labeled samples is needed. In the case of unbalanced labeled samples, the result is not quite accurate of the category which has the minimum quantity. We consider that among a large number of comments, many of the top comments are positive but without involving valuable opinion. For the order depends on the user's level and the number of praises, while some really valuable comments may sink, it's hard for people to find them. The main work of this paper is to classify the comments into four types: valuable, negative, useless and normal. This paper has chosen ST-SVMs model, which realizes multi-classification through using SVM three times. At the same time, we proposed a concept of "Topic Proximity" which used LDA to extract the topic then used word2vec to calculate the similarity between the comments and topic. Then we give a combined feature, which consist of vector for comments text, sentiment intensity and topic proximity as the input features of the ST-SVMs model.