Neighborhood Distance Estimation for Tree-Based Hybrid Genetic Algorithm for Density Based Data Clustering

Mashuk Khan, Md. Ahied Mahi Chowdhury, Noureen Alam Meem, Mozammel H. A. Khan · 2022

Clustering algorithms partition data points of a dataset depending on their similarity. As this process is unsu-pervised, validation is a crucial part of this method. Generally, the optimal clusters are verified using external information about the dataset which represents the true clusters of the data points. But in real life datasets, these ground truth information are not always present. That is why internal validation is used which can validate the clusters only using features of the data points. In this paper, internal validation (S_Dbw) is used on a previously proposed tree-based hybrid genetic algorithm for density based data clustering so that the minimum neighborhood distance can be estimated without using ground truth information. Four datasets from UCI Machine Learning Repository were used in this experiment. The proposed model outperforms the existing algorithms for Seeds and Parkinson's datasets. For Wine dataset the result is close but less accurate. But for Iris dataset, the model does not perform well because of overlapping of the clusters.

Read the paper · More papers on PaperTik