Text Categorization Based on a Similarity Approach

Cha Yang, Jun Wen · 2007

Text classification can efficiently enhance the text processing capability by automatically sorting out them according to defined collection of categories.This paper uses TFIDF method to represent documents, and set the NGramSize value to be 6.Word Frequency vector is used to measure and distinguish different features on documents.The Similarity Approach uses Cosine function to construct the classifier.The experiment results indicate that proposed algorithm yields good performance with the accuracy up to 98%.

Read the paper · More papers on PaperTik