Performance Enhancement of Hadoop MapReduce by Combining Data Inside the Mapper
Md. Niaj Shahriar Shishir, Mohammad Abu Yousuf · 2021 2nd International Conference on Robotics, Electrical and Signal Processing Techniques (ICREST) · 2021
Hadoop is a framework for storing and processing large scale data. Hadoop Distributed File System (HDFS) handles the storage part and MapReduce does the data processing while Yet Another Resource Negotiator(YARN) manages all the resources of the cluster. MapReduce is a programming framework for user defined Mapping and Reducing. Key-value pairs generated by Mapper are unfiltered and recurring, transferred to Reducer results in a huge unnecessary data throughput creating bandwidth overdraw. In this paper, a new algorithm called Inner Map Combining(IMC) for Map phase is proposed. The values of recurring keys get combined by this algorithm inside the Mapper. A successful testing was conducted in 12 different way to test the efficiency of the algorithm. The test shows that IMC implemented with Default Combiner(DC) of MapReduce is 65% more efficient than basic MapReduce program without any combiner. This work can be followed while MapReduce programming for significantly faster computation.