Adapting k-means for Clustering in Big Data
Mugdha Jain, Chakradhar Verma · International Journal of Computer Applications · 2014
Big data if used properly can bring huge benefits to the business, science and humanity.The various properties of big data like volume, velocity, variety, variation and veracity render the existing techniques of data analysis ineffective.Big data analysis needs fusion of techniques for data mining with those of machine learning.The k-means algorithm is one such algorithm which has presence in both the fields.This paper describes an approximate algorithm based on k-means.It is a novel method for big data analysis which is very fast, scalable and has high accuracy.It overcomes the drawback of k-means of uncertain number of iterations by fixing the number of iterations, without losing the precision.