Modular Data Clustering - Algorithm Design beyond MapReduce
Martin Hahmann, Dirk Habich, Wolfgang Lehner · 2014
In the context of Big Data, flexible and adjustable data analytics become more and more important, whereas an efficient, scalable and fault-tolerant execution is required as well. To fulfill the flexibility as well as the execution requirements, the specification of the analysis methods have to be in an appropriate and easy adjustable manner. The MapReduce approach has demonstrated that such flexible specification as well as scalable execution is possible and applicable. However, the MapReduce programming model is too generic and complicates the specification from a data analysis point of view. Therefore, we propose a novel programming approach using well-defined modular building blocks for a specific and highly utilized data analysis domain named data clustering in this paper. Our approach offers many advantages: (i) a unified and specific instruction set for data clustering which eases understanding and algorithm adaptation in an abstract way, and (ii) enables an efficient and scalable execution of all data clustering algorithms based on an efficient mapping of the unified instruction set to a specific target environment is possible.