A New Classification Algorithm for Large Scale of Chinese Texts
Hongwei Wang, Jianhui Wang, Yi Lei · 2006 IEEE International Conference on Service Operations and Logistics, and Informatics · 2006
Most of classifying methods are based on VSM in the current classification research, of which the widely-used method is kNN. But most of them are highly complicated on computation, and could hardly be used for classifying a large number of samples. Moreover, to them, the classifier must be rebuilt when adding or deleting the training samples, which make them poor in scalability. In this paper, two new concepts, mutual dependence and equivalent radius, are presented, based on which a new classifying method (called MDER) is offered. MDER can be used to classify a large number of samples and has good scalability. After a series of experiments of classifying Chinese documents, the conclusion are drawn that MDER outperforms kNN and CCC method, and can be used online to classify a large number of samples while keeping higher precision and recall