A Generalization of k-Means for Overlapping Clustering
Guillaume Cleuziou · 2007
This paper deals with overlapping clustering, a trade off between crisp and fuzzy clustering. It has been motivated by recent applications in various domains such as Information Retrieval or biology. We show that the problem of finding a suitable coverage of data by overlapping clusters is not a trivial task and we propose the algorithm OKM that generalizes the k-means algorithm combining a new objective criterion coupled with an optimization heuristic. Experimental results in the context of document clustering show that OKM first generates suitables overlaps between classes and then outperforms the overlapping clusters derived from fuzzy approaches (e.g. fuzzy-k-means). 1