Enhancing K-Means Clustering with Post-Redistribution

Aymen Takie Eddine Selmi, Mohamed Faouzi Zerarka, Abdelhakim Cheriet · Ingénierie des systèmes d information · 2024

Traditional K-means clustering may converge to suboptimal solutions due to local optima, impacting cluster balance and compactness.To fix this, we suggest an enhanced K-means algorithm that includes a new step for redistribution post-clustering that is based on the sum of squares errors (SSE) and diameter.Our approach introduces a redistribution step focusing on achieving balanced population distribution within clusters.Evaluation metrics include Davies-Bouldin Index (DBI) and Gini coefficient, quantifying improvements in cluster compactness and balance.We compare our method against traditional K-means on diverse datasets, such that a lower value indicates better clustering results.The post-clustering redistribution significantly reduces DBI and Gini coefficient, indicating enhanced cluster quality and balance.This improvement is consistent across various datasets, showcasing the method's reliability and generalizability.Our improved K-means algorithm achieves better cluster balance and compactness by redistributing post-clustering, which also reduces problems with local optima.The method's applicability extends to diverse domains, providing more reliable clustering outcomes with practical implications in areas such as customer segmentation, anomaly detection, pattern recognition, and resource optimization.

Read the paper · More papers on PaperTik