Analyzing & optimizing hadoop performance
Ankita Jain, Monika Choudhary · 2017
Processing and analyzing BigData in timely and cost effective manner is an important but tedious job. It is highly desired that any BigData management framework must process and analyze data within a fraction of seconds. There are various tools available for this purpose; one of such open source tool is Apache Hadoop. This tool is widely adopted for storage and management of BigData. Several methods have been suggested and implemented to analyze and enhance the Hadoop performance. This paper particularly focuses upon tuning configuration parameter approach. Although there are several configuration parameters which affect the performance of Hadoop, among that MapReduce related parameter has a significant impact. The objective of this work is to enhance the overall performance of Hadoop by reducing job execution time. Reduction in time is obtained by tuning some of MapReduce associated parameters. The right understanding of these parameters is very crucial since varying parameters with improper values can show a negative impact on overall performance. In this paper, proposed approach saves job execution time and optimizes disk usage efficiently. It significantly improves the overall performance of Hadoop by 38.51% over the base system in the heterogeneous environment.