Saksham: Resource Aware Block Rearrangement Algorithm for Load Balancing in Hadoop

Ankit Shah, Mamta Padole · Procedia Computer Science · 2020

Big Data Analytics demands huge processing power. Distributed computing seems to be a viable solution with the usage of commodity computers. But, processing big data optimally, on distributed systems poses the number of challenges, such as seamlessly integrating different resources of distributed systems. Hadoop provides this ease and hence is one of the most commonly used tools for the processing of big data. But, it is observed that the performance of Hadoop is compromised, when it comes to the heterogeneous environment. This paper proposes Saksham: The block rearrangement algorithm which optimizes the processing time by efficient file system management in the heterogeneous environment. Our proposed scheme successfully optimizes the big data processing over the homogeneous and heterogeneous environment. To achieve better performance for big data processing, we target two important aspects of heterogeneous distributed computing: file system management and process management. First, file system management basically controls the block placement and allows us to rearrange the blocks on specified nodes, considering the node’s storage and processing capability. Second, we use the concept of node labeling and scheduling to achieve better process management. The results prove that the proposed approach has optimized job execution time along with reduced latency and data-skew substantially.

Read the paper · More papers on PaperTik