Effective Reduction Of Network Traffic Cost In Map Reduce For Very Large Scale Data Applications
K. REDDAMMA · i-manager s Journal on Computer Science · 2017
In the recent time, the amount of data being generated is rapidly evolving from terabytes to many petabytes. The traditional data processing tools available, fail to handle the large data size. The tools are evolving to overcome the inadequacies existing and provides a solution to manage and process massive volumes of heterogeneous data. The current technologies support to manage large data storage, processing, analyzing, effective decision making, etc. When the size of the data increases the network traffic will also increase as well. The High Performance Computing (HPC) and Grid computing [1], [2] have been adopted for large scale data processing. The idea implemented in HPC is to work with the cluster of machines and access a shared distributed file system. This is especially used for highly computational intensive jobs. The issue in the above computing approach is network bandwidth. If the network traffic develops largely, network bandwidth is restricted, and hence computing nodes in the high performance cluster become idle. The HDFS (Hadoop Distributed File System) is a distributed file system constructed to run on commodity hardware. It has many features similar to other distributed computing. HDFS is designed to provide more fault-tolerance and to deploy on low-cost hardware. It is also designed to process and manage massive volumes of heterogeneous data. The Hadoop can store, analyze and process large volumes of data. Also, Hadoop has capacity of large system, faulttolerance, scalability, and availability. MapReduce [3] is a software framework used to write applications that process large amounts of data in parallel on clusters of commodity hardware. A MapReduce job first divides the data into individual chunks which are processed by Map jobs in parallel. The outputs of map sorted by the framework are then input to the reduce tasks. In general, the input and the output of the jobs are both stored in a file system. The MapReduce online is a modified version of Hadoop [4] MapReduce, which supports online aggregation and reduces response time. This approach has the advantage of simple recovery in the case of processing into two phases: (a) Map phase and (b) Reduce phase. A map uses a common partitioner records that is partitioned into