InfoSift: Adapting graph mining techniques for text classification
Manu Aery, Sharma Chakravarthy · 2005
Text classification is the problem of assigning pre-defined class labels to incoming, unclassified documents. The class labels are defined based on a set of examples of pre-classified documents used as a training corpus. Various machine learn-ing, information retrieval and probability based techniques have been proposed for text classification. In this paper we propose a novel, graph mining approach for text classifica-tion. Our approach is based onthe premise that representa-tive – common and recurring –structures/patterns can be ex-tracted from a pre-classified document class using graph min-ing techniques and the same can be used effectively for clas-sifying unknown documents. A number of factors that influ-ence representative structure extraction and classification are analyzed conceptually and validated experimentally. In our approach, the notion of inexact graph match is leveraged for deriving structures that provide coverage for characterizing class contents. Extensive experimentation validate the selec-tion of parameters and the effectiveness of our approach for text classification. We also compare the performance of our approach with the naive Bayesian classifier.