An Efficient Filtered Classifier for Classification of Unseen Test Data in Text Documents

G. Naga Chandrika, Edara Srinivasa Reddy · 2017

Rapid development of information technology has increased the availability of Electronic documents and the task of automatic classification of e-documents play an important role for organizing the information in large data repositories. In addition, many researchers proposed various algorithms for classification, but these approaches need to filter the data before classification. Keeping these limitations, we address the problem of classifier for classification of unseen data in text documents where document data distribution is not homogeneous. In this study, we used a Filtered Classifier on text documents that has passed through an arbitrary filter. To classify the text documents C4.5 classifier was used, the structure of the filter is based on training and testing data sets that are processed by the filter without changing the structural behavior. In addition, Fayyad and Irani's discretization method is used as a preprocessing that discretize a range of numerical attributes in the text document data set into nominal attributes. For classification, we use C4.5 decision tree classifier. Four datasets such as CNAE-9, 20 Newsgroups, Twitter and Reuter-21578 were employed to test the unseen test documents and test the efficiency of the Filtered Classifier. Experimentation is described in detail and the results show improved classifier accuracy for classification.

Read the paper · More papers on PaperTik