Document Classification using Recommendation Keywords and Machine Learning
Moon-Hyeon Kim, Yeonghyeon Gu, Seong-Joon Yoo · Jeongbo gwahaghoe nonmunji. so'peuteuweeo mich eung'yong · 2011
We propose a novel approach for the automatic classification of Korean text documents containing product recommendations using machine learning and rules. Most of Korean product recommendations include comparative keywords such as 'than', or recommendation keywords including 'recommend', 'superior', 'excellent', and 'overwhelming victory'. We apply some rules or machine learning based classifier to select candidate sentences including such keywords and sort out only the recommendation sentences. The result of classifying 1,336 documents, including five comparative and recommendation keywords using Naive Bayes and Bayesian Net shows a recall rate of 88.3% and a precision of 83.5%. In the future, hopefully, there will be further studies on approaches to classification of generalized recommendation sentences in terms of more comparative and recommendation keywords. The idea of our previous work on mining comparative only sentences published in CSA2009 can be exploited in classifying recommendation sentences by adding the features proposed in this paper.