streamingRPHash: Random Projection Clustering of High-Dimensional Data in a MapReduce Framework
Jacob Franklin, Sam Wenke, Sadiq Quasem, Lee A. Carraher, Philip A. Wilsey · 2016
The expanding needs for analysis on large datasets has increased as the amount and availability of data continues to grow. The size and format of this data makes manual analysis infeasible and has motivated the drive for automated methods such as data clustering. Among the most commonly used clustering algorithms, K-means has been proven as one of the most popular choice that delivers acceptable results in reasonable time. For many years, K-means has proven to be statistically efficient and easy to implement. While K-means is widely used for clustering streaming data, it has performance issues when it comes to robustness with noise, parallelism and working with very large, high data sets. In particular, K-Means (and other conventional techniques) for data clustering do not parallelize or scale well with the increasing dimensionality of data.