An Approach to Improve Classification Accuracy in Very Large Datasets
Marilyn G. Kletke, Dursun Delen, Jin-Hwa Kim · Journal of the Association for Information Systems · 2004
In this paper we present a study that suggests a two-step approach, called the Iterative Refinement Algorithm (IRA), for improving the classification accuracy of inductive learning algorithms applied to very large datasets.We present the preliminary test results for IRA compared to other prediction methods including logistic regression, discriminant analysis, neural networks, C5, CART, and CHAID on a census dataset of approximately five million records.We offer IRA with the belief that it is an incremental step towards overcoming the limitations of current data mining tools as they are applied to today's massive datasets.