Implementation of Distributed Searching and Sorting using Hadoop MapReduce
Sanjeev Kumar Pippal, Ankur Shukla, Dharmender Singh Kushwaha · 2014
This paper focuses on implementation of MapReduce programming model on Hadoop cluster for parallel processing of huge amount of data efficiently. There is deluge of data everywhere and we need to process these data efficiently to take decisions and to reach at some meaningful conclusion so that the information can be made useful. Processing large data set is a big problem as most of the data in unstructured. So we can't apply any algorithms since efficiency is a issue. In that case we need to process complete data set for almost every query on the data to get the desired result. If we process this huge amount of data on a single machine in linear fashion then we face problems like data storage issues and response time. So we need a mechanism which can process huge data sets in parallel fashion on distributed machines. MapReduce programming model is one such widely accepted model for parallel processing of data. While executing queries on the file containing 10 million records, it is observed that as the number of data nodes in the Hadoop cluster is increased, the minimum performance improvement obtained is over 42%. When implementing Distributed Grep Search vs Normal Implementation of Grep for a input file that has 10 million records containing user name, operation, ip address, with 3 data nodes, an improvement of over 18% is observed. For word count application, a substantial improvement of over 200% is observed with 3 or more data nodes. With our input size secondary sort outperforms insertion sort and bubble sort algorithms.