The Study and Application of Hadoop across Multiple Clusters

Sun Shengtao, Aizhi Wu, Xiao-Yang Liu · Advances in engineering research/Advances in Engineering Research · 2014

Hadoop is a widely applied tool for large-scale data-intensive computing in big data, but it can only be implemented on single cluster environment.In this paper, we focus on the application of Hadoop across multiple clusters and dedicate to solve the key problems of data sharing and task scheduling among clusters.A hierarchical distributed computing architecture of Hadoop across multiple clusters is designed.The virtual HDFS and job adapter are proposed to provide global data view and task allocation across multiple data centers.The job submitted by user to this platform is decomposed automatically into several sub-jobs and then allocated to corresponding cluster by location-aware manner.A prototype based on this architecture is presented and currently applied in the distributed spatial information processing across spatial data centers.

Read the paper · More papers on PaperTik