Research on text categorization based on LDA
Peng Cheng · Computer Engineering and Applications Journal · 2011
When the text corpuses are high-dimensional and large-scale,the traditional dimension reduction algorithms will expose their limitations.A Chinese text categorization algorithm based on LDA is presented.In the discriminative frame of Support Vector Machine(SVM),Latent Dirichlet Allocation(LDA) is used to give a generative probabilistic model for the text corpus,which reduces each document to fixed valued features——The probabilistic distribution on a set of latent topics.Gibbs sampling is used for parameter estimation.In the process of modeling the corpus,a latent topics-document matrix associated with the corpus has been constructed for training SVM.Standard method of Bayes is used for reference to get the best number of topics.Compared to Vector Space Model(VSM) for text expression combined SVM and the classifier based on Latent Semantic Indexing(LSI) combined SVM,the experimental result shows that the proposed method for text categorization is practicable and effective.