Parallel Approach in Data Mining Based on Hadoop Cloud Platform

Baoyuan Qi · Jisuanji fangzhen · 2013

The cloud platform has been dealt in industry with large-scale high-dimensional data.A variety of heterogeneous systems have been simulated as one system,in which data mining on Hadoop will encounter the issues,such as the globalization of data models,the random write operations of HDFS files,and the duration of data life.For practical large-scale high-dimensional data mining,an efficient data mining framework on Hadoop was proposed to solve these problems,which used databases to simulate the linked list structure,and provided a distributed algorithm for structures of tree and graph model.Based on it,a statistical algorithm-Yscore binning – was proposed,as well as the DB-tree and KD-tree building algorithm.The Vega cloud was used as a simulation of Hadoop cluster.The experimental data shows that the framework and the algorithm is practical and feasible,and may be expanded to other areas outside of data mining.

Read the paper · More papers on PaperTik