An efficient technique to improve resources utilization for hadoop MapReduce in heterogeneous system
Ahmed Qasim Mohammed, Rajesh D. Bharati · 2017
Oughties witness releasing one of the most reputed platform for processing and storing BigData which known by a strange name is Hadoop, mainly Hadoop consist of two main application MapReduce for processing data and Hadoop Distributed File System (HDFS) for storing data, in fact even all the good features of Hadoop there's still some struggles specially with high speed of data growth, therefor there is need to improve the performance of the main two components to increase Hadoop capability to treat data in efficient way. Main focus is improving resources utilization in MapReduce as a result of this there will be maximum usage of resources and minimizing time for processing data so we implemented different techniques on different level of MapReduce, our work started by adding a classifier level by using lightweight classification algorithm Support Vector Machine (SVM) to overcome heterogeneity issues that face Hadoop and generate problem to assign proper job to proper slave node, by this technique we decreased failed Tasks to approximately zero. Second allocating slot dynamically according to needs by editing allocation rules to permit slot to process either Map task or Reduce task on same slot according to the needs of user. Third, technique is finding a balance for the tradeoff between a single job and batch of job by guessing the performance of executing dynamically. Fourth, proposing a technique to improve data locality without impacting fairness by using pre schedule slot. In overall using improving on different levels show improvement in performance of MapReduce with increasing data locality without effecting fairness and decreased running time by 30% depending on variation of Jobs and there requirements, also improve processing jobs on heterogynous cluster with minimizing job failed.