A Comparison Study of Similarity Measures in Rough Sets Clustering

Arnold Szederjesi-Dragomir, Radu Găceanu, Horia F. Pop, Costel Sârbu · 2019

The proper selection of the similarity measure may be of paramount importance for the clustering result especially if the clusters overlap. Therefore, the aim of this paper is to study the influence of some widely known similarity measures on the clustering structure. The analysis uses a clustering algorithm based on rough sets which is able to discover hybrid data (outliers and instances that are close enough to more than one cluster). In this paper, we examine the similarity measures influence on the overlapping regions as well. Since the standard data sets do not offer specific information regarding the overlapping areas, we also propose an approach for extracting this data in order to be used for benchmarking purposes. Experiments conducted on standard data sets outline the importance of the proper similarity measure selection.

Read the paper · More papers on PaperTik