A kind of Prefetching Data Way to Hadoop MapReduce Environments
Hui Xia, Peng Wu · 2016
In MapReduce environment, problem of data redundancy,a large number of tasks to be processing and mass data storage come up, in order to solve these problems, we put forward the way of data prefetching , preprocessing of hid the remote data access latency, by adjusting the allocation of resources to Reduce the business, we put forward the method of than before, caused by the Reduce tasks of remote data access performance problems caused by time delay and resource competition system, the hidden by prefetching data the method Reduce task of remote data access latency, and Reduce task control through the resource allocation, to Reduce resource competition caused the Reduce task, the experiment results show that with the default Hadoop graphs and Hadoop Online Prototype (HOP), compared with the method the system performance can be improved by more than 10%. Reduce Performance issues caused by the taskMapReduce model based on the preparation of the job contains a large number of Map tasks and Reduce tasks, Map task processing operations of the original input data, resulting in form of records, these records as intermediate results stored in the local node; Reduce tasks to deal with these Record, resulting in the final output of the job, but for each Reduce task, only the record containing the specific key in the intermediate result is processed.During the execution, the Reduce task first copies the records corresponding to the specific keys from all the nodes, then sorts the records, and then performs the specified operations on the sorted records.As the number of Map tasks is very large and distributed among different nodes, Reduce tasks to complete the record copy, you must perform a large number of remote I / O operations; these operations will be in the process of introducing a large number of remote data access latency.