Comparison of Centroid-Based Clustering Model Performance on Categorical Dataset

Fitri Nuraeni, Abdul Syukur, Aris Marjuni, Nova Rijati, Dede Kurniadi · 2024

Clustering is a popular data grouping technique in statistical analysis, especially for unlabeled datasets. Centroid-based clustering algorithms such as k-means, fuzzy cmeans, and k -medoids are generally used for numerical data but have limitations for categorical data. This research compares the performance of various centroid-based clustering models on categorical datasets. This research method involves collecting ten categorical datasets from the UCI Machine Learning Repository and testing various categorical to numeric feature conversion techniques such as One Hot Encoding, Bloom Filter Encoding, and Wide and Deep Learning. The clustering models tested include k-means, Fuzzy C-means, k-medoids, K-Modes (Huang), K-Modes (Cao), Fuzzy K-Modes, and Like K-Means. The evaluation used the Silhouette Score, Cluster Homogeneity, Rand Index, Fowlkes-Mallows Index, and Purity metrics. The research results show that K-Modes (Cao) and K-Modes (Huang) show competitive performance. OHE+FCM, WDL+Kmed, and K-Modes (Cao) were the most effective clustering model for categorical datasets. This research shows the importance of feature conversion techniques and selecting a suitable clustering model to improve the quality of categorical data grouping.

Read the paper · More papers on PaperTik