Analysing the effect of different Distance Measures in K-means Clustering Algorithm

Trushali Jambudi, Savita R. Gandhi · GLS KALP: Journal of Multidisciplinary Studies. · 2024

Distance metrics are primary means for measuring the distance between two objects and used as the principal means of deciding the similarity or dissimilarity between the data to be clustered. Different distance measures are applied by different clustering algorithms for the purpose of grouping objects into clusters. The use of a particular distance metric can affect to a great extent the performance of a clustering algorithm and hence the outcome also. In this paper, we analyze the impact of various distance measures in the performance of K-means algorithm. We first describe the different distance measures that are commonly used with K-means algorithm, followed by application of the K-means clustering algorithm with each of these distance measures on various synthetic and real standard clustering data sets. To measure the impact of each distance measure on the performance of K-means algorithm, we have deployed various performance evaluation metrics.

Read the paper · More papers on PaperTik