GSHetero - Grouping and Heterogeneity-Aware Data Placement to Improve MapReduce Performance in Hadoop
S. Vengadeswaran, P. Dhavakumar, Viji Viswanathan · IEEE Networking Letters · 2025
The execution of MapReduce (MR) applications in Hadoop cluster poses significant challenges due to the non consideration of 1. Grouping semantics in Data-intensive applications, 2. Heterogeneity in the computing nodes resulting in suboptimal block distribution, concentrating execution on fewer nodes, thereby increasing processing time and reducing data locality. This letter proposes improved data placement by exploiting grouping semantics and heterogeneity (GSHetero) to boost MR performance. Initially, the execution traces will be analyzed to identify the data access pattern. The grouping semantics are extracted by applying the MCL algorithm. Then GSHetero algorithm is proposed which re-organises the default data layouts based on grouping semantics to ensure higher parallelism. The efficiency of the GSHetero is demonstrated by the 10-node Hadoop cluster deployed on the cloud by executing the Linear Regression over the weather dataset. The results show that GSHetero improves data locality by 27.4% and CPU utilization by 47%. The efficiency of the GSHetero is also demonstrated by executing Hadoop benchmark (WordCount) on varying cluster sizes (15, 20 nodes) for varying workloads.