Implementation of Aggregation of Map and Reduce Function for Performance Improvisation
Varsha B.Bobade · Zenodo (CERN European Organization for Nuclear Research) · 2016
Big Data is term that refers to data sets whose size (volume), complexity (variability), and rate of growth (velocity) make them difficult to capture, manage, process or analyzed. To analyze this enormous amount of data Hadoop can be used. Hadoop is an open source software project that enables the distributed processing of large data sets across clusters of commodity servers. I proposed a modified MapReduce architecture that allows data to be pipelined between operators. This reduces completion times and improve system utilization for batch jobs as well. I present a modified version of the Hadoop MapReduce framework that supports online aggregation, which allows users to see early returns from a job as it is being computed. The objective of the proposed technique is to signicantly improve the performance of Hadoop MapReduce for efficient Big Data processing.