Combining Distributed and Multi-core Programming Techniques to Increase the Performance of K-Means Algorithm
Ilias Κ. Savvas, Dimitrios C. Tselios · 2017
The last years, huge masses of data are produced or extracted by computational systems and independent electronic devices. To exploit this resource, novel methods must be employed or the established ones may be altered in order to confront the issues that arise. One of the most fruitful techniques, in order to locate and use information from data sources is clustering, and k-means is a successful representative algorithm which clusters data according specific characteristics. However, its main disadvantage is the computational complexity which proves the techniques very unproductive to apply on big datasets. Although k-means is a very well studied technique, a fully operational distributed version combining the multi-core power of today machines, have not been accepted yet by the scientific community. In this work, a three phase distributed/multi-core version of k-means is presented. The obtained experimental results are very promising and prove the correctness, the scalability, and the effectiveness of the proposed technique.