The Research on Distributed Adaptive Text Classification

Xiao-Gao Yu · 2008

The automated categorization of documents into predefined labels has received an ever-increased attention for the exponential growth of documents on the Internet and the emergent need to organize them in the recent years. K-nearest neighbors is a widely used classifier in text categorization community because of its simplicity and efficiency among all these classifiers. However, K-nearest neighbor classification (KNNC) still suffers from inductive biases or model misfits that result from its assumptions, such as the presumption that training data are evenly distributed among all categories. In this paper, a new refinement strategy (DBKNNC) for the KNN classifier is proposed, which adopts sum-of-squared-error criterion to adaptively select the contributing part from these neighbors and classifies the input document in term of the disturbance degree which it brings to the kernel densities of these selected neighbors. DBKNNC is not sensitive to the parameter k and achieves significant classification performance improvement on imbalanced corpora according to the experimental results.

Read the paper · More papers on PaperTik