Dynamically Controlling Node-Level Parallelism in Hadoop
Kamal Kc, Vincent W. Freeh · 2015
Hadoop is a widely used large scale data processing framework. Applications run in Hadoop as containers, the concurrency of which affects completion time of an application as well as system resource usage. When there are too many concurrent containers, resource bottlenecks occur and when there too few, system resources are underutilized. The default and best practice settings underutilize resources which results in longer application completion times. In this work, we develop an approach to dynamically change the parallelism for concurrent containers to suit an application. Our approach ensures efficient utilization of resources and avoids bottlenecks for all types of MapReduce applications. Our approach improves performance of MapReduce applications by as much as 28% and 60% respectively when compared to the best practice and default settings.