The Research of Chinese Short-text Classification Based on Domain Keyword Set Extension and HowNet

Xiangdong Li, Fan Gao, Cong Ding · 2016

To implement feature extension of short text and improve short text classification performance, this paper extracts the high frequency words and topic core words of each class of the training set as domain keyword set based on two different feature granularity, which are keyword and latent topic, and derives the topic probability distribution of the test text using LDA model, while some topic probability is greater than a certain threshold, extends the keywords of the topic into the testing text.Calculate the semantic similarity of the test text and the domain keyword set for each category by using HowNet.Experimental results show that the method proposed in this paper can effectively improve the short-text classification performance.

Read the paper · More papers on PaperTik