Novel technique for prediction analysis using normalization for an improvement in K-means clustering

Shruti Gupta, Abha Thakral, Shilpi Burman Sharma · 2016

Clustering is the unsupervised classification of patterns in a dataset. Clustering is widely used to discover distributed patterns and classify them as clusters. Clustering algorithms uses a similarity measure based on distance. In order to cluster data points, k-means uses Euclidean distance measure and central point choice. In the K-means clustering, data points will be stacked and a central point is chosen. From the central point chosen, Euclidean distance will be computed and on that basis clusters will be assigned to the data points. One of the drawbacks of K-means is that numbers of clusters has to be provided due to which some data points remains un-clustered. In this paper, we propose a clustering calculation through which number of clusters can be characterised naturally. The proposed technique will improve accuracy and decrease clustering time moreover cluster quality will also be improved through multiple iterations.

Read the paper · More papers on PaperTik