Application of Genetic Algorithms to Data Mining
Robert E. Marmelstein, Wright-Patterson Afb · 1997
Data Mining is the automatic search for interesting and useful relationships between attributes in databases. One major obstacle to effective Data Mining is the size and complexity of the target database, both in terms of the feature set and the sample set. While many Machine Learning algorithms have been applied to Data Mining applications. There has been particular interest in the use of Genetic Algorithms (GAs) for this purpose due to their success in large scale search and optimization problems. This par per explores how GAs are being used to improve the performance of Data Mining clustering and classification algorithms and examines strategies for improving these approaches. What is Data Mining? Data Mining is an umbrella term used to describe the search for useful information in large databases that cannot readily be obtained by standard query mechanisms. In many cases, data is mined to learn information that is unknown or unexpected (and therefore interesting). To this end, Data Mining applications typically have as their goals one or more of the following: Finding Patterns: Patterns are clusters of samples that are related in some way relative to the database attributes. Discovered patterns can be unconstrained ("Identify all clusters consist of at least 10 percent of the population") or directed ("Find population clusters that have contracted cancer"). Deriving Rules: Rules describe associations between attributes. Accordingly, clusters provide a good source for rules. Rules are typically defined in terms of IF/THEN statements. An example rule might be: IF ((Age> 60)AND (Smoker = TRUE)) THEN Risk(Cancer) = 0.7;