Adaptive Node-Oriented Data Placement for Heterogeneous Hadoop Clusters

Vishnu Prasad Verma, Santosh Kumar, Nenavath Srinivas Naik · IEEE Access · 2025

Modern society is experiencing a data explosion thanks to rapid IT development and the increasing intelligence of devices. The vast and complex data can be utilized to extract actionable insight using a big data processing framework. Hadoop is a popular big data processing framework on heterogeneous commodity hardware. While Hadoop offers a robust framework for large-scale, data-intensive tasks via its MapReduce paradigm, hardware heterogeneity across nodes often leads to straggler effects that degrade Hadoop cluster performance. This paper introduces the Adaptive Node-Oriented Data placement for Efficient Hadoop Execution (ANODE) method, which leverages historical job execution data to dynamically assess each node’s processing capability. By employing an agent-based mechanism, ANODE optimizes block allocation within the data node, alleviating imbalances caused by Hadoop’s default uniform placement strategy. Experimental results on a heterogeneous eleven-node Hadoop cluster demonstrate that ANODE reduces job completion times by up to 25%, significantly enhancing data locality and resource utilization compared to the default approach.

Read the paper · More papers on PaperTik