Impact of Data Transfer to Hadoop Job Performance
Hyuma Watanabe, Masatoshi Kawarasaki · 2014
Hadoop is a distributed processing platform for analyzing a large amount of data. It uses MapReduce framework for distributed parallel processing where reduce computation uses all the relevant map computation results obtained in other nodes, thus needs massive amounts of data transfer between nodes. This paper explores how such interdependence between nodes affects the job performance in Hadoop cluster and clarifies the mechanism of job performance deterioration. For this purpose, we built two kind of experimental Hadoop clusters using real machines in our laboratory and virtual machines on Amazon EC2 and tracked the progress of tasks which proceeds in parallel. As a result, we revealed that delay in one task caused by congestions in disk I/O or data transfer propagates to other tasks and deteriorates the overall job performance significantly. Furthermore, we found that speculative task execution brings adverse effects when the task delay is caused by disk I/O.