Approximation to the K-Means Clustering Algorithm using PCA

Sathyendranath Malli, H. R. Nagesh, B. Dinesh Rao · International Journal of Computer Applications · 2020

Healthcare is an emerging domain that produces data exponentially.These massive data contain a wide variety of fields, which lead to a problem in analyzing the information.Clustering is a popular method for analyzing data.Data is split into smaller clusters having similar properties and is then analyzed.The K-Means algorithm [1] is a well-known technique among clustering methods.In this paper, an efficient approximation to the K-means problem targeted for large data by reducing the number of features to one through Principle Component Analysis(PCA) is introduced.This data is clustered in one dimension using the K -means algorithm.Intra-cluster RMS error in the modified algorithm is compared with the K-means algorithm in m dimensions and is found to be reasonable.The time taken by the modified algorithm is significantly less when compared to the K -means algorithm.

Read the paper · More papers on PaperTik