A distributed-population genetic algorithm for discovering interesting prediction rules
Edgar Noda, Alex Alves Freitas, Akebo Yamakami · 2002
In data mining the quality of prediction rules basically involves three criteria: accuracy, comprehensible and interestingness. The majority of the rule induction literature focuses on discovering accurate, comprehensible rules. In this paper we also take these two criteria into account, but we go beyond them in the sense that we aim at discovering rules that are interesting (surprising) for the user. The search is performed by a distributed genetic algorithm (DGA) specifically designed for the discovery of interesting rules. DGAs constitute an interesting approach to tackle the premature convergence problem in evolutionary algorithms. In our approach the partition of the search space in semi-isolated subpopulations (demes) represents a subdivision of the task. We model the migration procedure of DGAs as an explicit means to promote cooperation among the demes. The algorithm addresses the dependence modeling task of data mining, where different rules can predict different goal attributes. This task can be regarded as a generalization of the very well known classification task, where all rules predict the same goal attribute. This paper also compares the results of the DGA with the results of a single population genetic algorithm to discover interesting rules. 1