Dynamic Colocation Algorithm for Hadoop
B Ganesh Babu, T.P. Shabeera, S. D. Madhu Kumar · 2014
Hadoop is a widely accepted platform for developing large-scale data intensive applications. It is an open source implementation of Google's MapReduce framework. The current data placement policy of Hadoop distributes the data among DataNodes using random placement policy for simplicity and load balance. This simple data placement is good for Hadoop applications that used to access data from a single file. But if any application needs data from different files simultaneously, the performance normally degrades. Identifying the related files and placing them in the same DataNode or in adjacent DataNodes reduces network overhead and reduces the query span. We propose a Dynamic Colocation Algorithm, where the average number of machines that are involved in processing a query decreases by colocating the datasets, that are frequently accessed together and hence reduces the network overhead. Our technique checks the relations between datasets dynamically and rearrange the datasets according to their relations. Our experimental results show that, after colocation there is a significant reduction on the execution time of MapReduce programs.