Clustering by a genetic algorithm with biased mutation operator

Benjamin Auffarth · 2010

In this paper we propose a genetic algorithm that partitions data into a given number of clusters. The algorithm can use any cluster validity function as fitness function. Cluster validity is used as a criterion for cross-over operations. The cluster assignment for each point is accompanied by a temperature and points with low confidence are preferentially mutated. We present results applying this genetic algorithm to several UCI machine learning data sets and using several objective cluster validity functions for optimization. It is shown that given an appropriate criterion function, the algorithm is able to converge on good cluster partitions within few generations. Our main contributions are: 1. to present a genetic algorithm that is fast and able to converge on meaningful clusters for real-world data sets, 2. to define and compare several cluster validity criteria.

Read the paper · More papers on PaperTik