An Equilibrium Approach to Clustering: Surpassing Fuzzy $C$-Means on Imbalanced Data

Yudong He · IEEE Transactions on Fuzzy Systems · 2025

Most fuzzy clustering algorithms are based on the well-known Bezdek’s fuzzy$C$-means (FCM). However, FCM fails when the data is imbalanced (class sizes are highly unequal). This issue arises because in FCM all data points have only attraction to each cluster prototype (i.e., centroid), causing centroids to be biased toward the large class with the most data points. This article proposes a novel equilibrium$K$-means (EKM) for imbalanced data, where data points exert both attraction and repulsion on centroids. The equilibrium between opposing forces reduces the learning bias toward large clusters, leading to more meaningful clusters on imbalanced data. Unlike FCM, EKM explicitly models the relationship between membership and centroid via equality constraints, avoiding the pitfalls of uniform effect. We derive closed-form centroid update equations proven to converge exponentially fast. Experiments are conducted on four artificial and 16 real-world datasets. The results demonstrate that EKM significantly outperforms 13 State-of-the-Art methods on imbalanced datasets while maintaining competitive performance on balanced data. EKM achieves an average improvement of 0.22 in normalized mutual information, 0.31 in adjusted rand index, and 0.21 in clustering accuracy over FCM on 10 real-world imbalanced datasets, with comparable computational efficiency in theory and practice.

Read the paper · More papers on PaperTik