Research of English Text Classification Methods Based on Semantic Meaning

Lin Lv, Yushu Liu · 2006

To overcome the limitations of traditional text classification approaches based on bag-of-words representation and to effectively incorporate linguistic knowledge and conceptual index into text vector space representation, based on WordNet thesaurus and latent semantic indexing (LSI) model, combinative method of them is presented to realize naive Bayes text classification and simple vector distance text classification, and five groups of contrastive experiments are made respectively. The results show that the accuracy rates of the two text classification methods are both gradually advanced along with more and more in-depth semantic analysis, which indicates that semantic mining is very important and necessary to text classification. The comparative analysis of the related work is also given.

Read the paper · More papers on PaperTik