The Capability Analysis on the Characteristic Selection Algorithm of Text Categorization Based on F1 Measure Value
Shaojun He, Jin Cao, Ruixu Guo, Wang Guijun · 2012
The text categorization is an important aspect in the processing of nature languages. It can be used to identify the categorization information within the nature languages, consequently, the clutter problem, directional detection and scout of information has been solved. The general processing of text categorization is proposed in this paper. Taken Sogou datasets as the target, the capability of several typical characteristic selection algorithms have been analyzed in KNN classification machine with different characteristic dimensions and classification methods, while the text categorization experiment is based on F1 measure value.