An Approach to Enhance the Performance of Hadoop MapReduce Framework for Big Data
Subhash Chandra, Deepak Motwani · 2016
Data analysis is becoming one of the highest research topic among researchers. Information is the baseline of every small and big organization. Everyone wants relevant information for their business to grow faster and bigger. Every organization wants to know what their customers like and dislike. This desirable information requires analysis of very large information stored in various places in different format. Hadoop MapReduce framework becoming a popular platform for processing so large amount of data in very efficient manner. It is used by organizations to process their customers information data sets. Hadoop process datasets in distributed parallel processes by using its HDFS and MapReduce model. Hadoop optimization is requiring more attention from researchers and programmers. Many approaches is already developed to make Hadoop framework optimized. These approaches includes performances tuning and efficient clustering formation. In this research work we have developed Optimal Approach to Improve the Performance of Hadoop framework. K-Means and K-Medoids are well known clustering approaches for clustering inside Hadoop. In proposed approach a modified K-Medoids clustering algorithm has been developed which gives better result for processing inside Hadoop. The research work is tested inside multi node Hadoop environment.