A performance analysis of MapReduce applications on big data in cloud based Hadoop

Parth Gohil, Dweepna Garg, Bakul Panchal · 2014

MapReduce is one of the most popular programming model for big data analysis in Distributed and Parallel Computing Environment. It is used for implementing parallel applications. With the growing development of mobile Internet and cloud computing, the issues related to big data have been a matter of concern in both industry and academy. There are several platforms for users to develop their applications based on MapReduce framework such as Hadoop. Hadoop is a free, Java-based programming framework that supports the processing of large data sets in a distributed computing environment. This paper discusses various MapReduce applications like Wordcount, Pi, TeraSort, Grep in Cloud based Hadoop. We have shown experimental results of these applications on Amazon EC2 using two types of Ubuntu instances. In this paper, performance of above application has been shown with respect to execution time and number of nodes. We find in our research study that as the number of nodes increases the execution time decreases and performance increases.

Read the paper · More papers on PaperTik