Performance Research on MapReduce Programming Model
Ming Li, Guanghui Xu, Lifa Wu, Ji Yao · 2011
Map Reduce programming model is designed to process large data sets in-parallel on large clusers. But most organizations can't afford to built a large cluster, so building a small cluster to improve the efficience of time-consuming applications is a perfect solution. Besides, non-data-intensive programs are common. Is Map Reduce suitable for this kind of programs? In this paper, a small cluster consisting of 5 PCs is built, with its configuration adjusted. Then a distributed FTP scan program is written to test whether Map Reduce is suitable for small data sets, network-intensive program. Finally a distributed string search program is written to test the performance of Map Reduce on large data sets. The results show that Map Reduce can run efficiently on small cluster, and it's also suitable for small data sets, network-I/O-intensive programs.