Chinese Short Text Classification Based on Domain Knowledge
Xiao Xin Feng, Yang Shen, Chengyong Liu, Wei Liang, Shuwu Zhang · International Joint Conference on Natural Language Processing · 2013
People are generating more and more short texts. There is an urgent demand to classify short texts into different domains. Due to the shortness and sparseness of short texts, con-ventional methods based on Vector Space Model (VSM) have limitations. To tackle the data scarcity problem, we propose a new mod- el to directly measure the correlation between a short text instance and a domain instead of representing short texts as vectors of weights. We firstly draw domain knowledge for each user-defined domain using an external corpus of longer documents. Secondly, the correlation is calculated by measuring the proportion of the overlapping part of the instance and the domain knowledge. Finally, if the correlation is greater than a threshold, the instance will be classified into the domain. Experimental results show that the classifier based on the proposed model outperforms the state-of-the-art baselines based on VSM.