Genetic Wrappers for Constructive Induction in High-Performance Data Mining

William H. Hsu, Michael E. Welge, Thomas C. Redman, David Clutter · 2000

We present an application of genetic algorithm-based design to configuration of high-level optimization systems, or wrappers, for relevance determination and constructive induction. Our system combines genetic wrappers with elicited knowledge on attribute relevance and synthesis. We discuss decision support issues in a large-scale commercial data mining project (cost prediction for multiple automobile insurance markets), and report experiments using D2K, a Java-based visual programming system for data mining and information visualization, and several commercial and research tools. Our GA system, Jenesis [HWRC00], is deployed on several network-of-workstation systems (Beowulf clusters). It achieves a linear speedup, due to a high degree of task parallelism, and improved test set accuracy, compared to decision tree learning with only constructive induction and state-space search-based wrappers [KJ97]. 1 GENETIC WRAPPERS AND KDD Our commercial decision support project applies a large demographic and historical customer database from the Allstate One Company project to a predictive classification problem in automobile insurance underwriting. The data mining (DM) pipeline we have developed comprises descriptive statistics, interactive visualization, data aggregation, data clustering, relevance determination, and supervised inductive learning for classification [HWRC00]. The research presented here focuses on genetic optimization of the inductive learning steps, where relevance determination is the objective of our attribute subset selection and synthesis (aka feature selection, extraction, and construction) stage. For this purpose, we have developed a genetic algorithm that calibrates hyperparameters in a validation-based model selection wrapper [RPG+97, HWRC00] written in a D2K-based GA, Jenesis [HWRC00]. 2

Read the paper · More papers on PaperTik