Using latent Dirichlet allocation to improve text classification performance of support vector machine

Yaw-Huei Chen, Shu-Fong Li · 2016

Text classification is an important task in natural language processing that aims to determine the category of a document. In the simplest settings, we adopt the bag-of-words model and convert documents in the corpus into term frequency vectors so that the classifier can process them. Because the bag-of-words model retains only the number of occurrences of each individual term, the classifier cannot use other syntactic and semantic information, which may lead to inaccurate classification results. In order to improve the accuracy of text classification, we use latent Dirichlet allocation (LDA) to extract the topic information so that we can add topic related features into the feature set representing the document. We explore different forms of term frequency and topic information and treat them as features for the support vector machine (SVM). The experimental results on three corpora indicate that the combined features can enhance the text classification accuracy.

Read the paper · More papers on PaperTik