An Automatic Clustering Algorithm Using NSGA-II with Gene Rearrangement

Hongchun Qu, Li Yin · 2020

As an unsupervised learning method in machine learning, clustering is a way to understand and learn from data. One of the most widely used clustering algorithms is the prototype-based clustering method, which needs to know the clusters number in advance. However, the number of clusters is often unlikely to be specified with prior experience. In this paper, we proposed a data clustering method, GR-NSGAII, which does not need to pre-set the number of clusters. Instead, the number of cluster and the results are generated automatically. Unlike the benchmark partitioning algorithm-based methods, it is a multi-objective optimization clustering approach based on genetic algorithm NSGA-II, in which only the objective functions were used to guide the clustering partitioning. Two optimization criteria, i.e., the sum of generalized sample variance and the Calinski-Harabasz index, were simultaneously optimized to control the number of clusters in a reasonable range. In addition, a technique named gene rearrangement combined with cluster merging was used to get better clustering results. In contrast with the state-of-the-art multi-objective evolutionary clustering methods, GR-NSGAII has better performance.

Read the paper · More papers on PaperTik