Improving Text Categorization By Using A Topic Model
Wongkot Sriurai · Advanced Computing An International Journal · 2011
Most text categorization algorithms represent a document collection as a Bag of Words (BOW).The BOW representation is unable to recognize synonyms from a given term set and unable to recognize semantic relationships between terms.In this paper, we apply the topic-model approach to cluster the words into a set of topics.Words assigned into the same topic are semantically related.Our main goal is to compare between the feature processing techniques of BOW and the topic model.We also apply and compare between two feature selection techniques: Information Gain (IG) and Chi Squared (CHI).Three text categorization algorithms: Naïve Bayes (NB), Support Vector Machines (SVM) and Decision tree, are used for evaluation.The experimental results showed that the topic-model approach for representing the documents yielded the best performance based on F 1 measure equal to 79% under the SVM algorithm with the IG feature selection technique.