Survey on improved Autoscaling in Hadoop into cloud environments

Masoumeh Rezaei Jam, Leyli Mohammad Khanli, Mohammad Kazem Akbari, Elham Hormozi, Morteza Sargolzaei Javan · 2013

Nowadays technologies for analyzing big data are evolving rapidly. Because of that models and methods to design and analyze parallel processing of data is done automatically. So MapReduce is one of these methods in order to overcome the complexity of very large data. MapReduce-based systems are suited for performing analysis at this scale since they were designed from the beginning to scale to thousands of nodes in a shared-nothing architecture. This model has been developed under a cloud computing platform. There are many implementations of MapReduce. One of them is the Apache Hadoop project that is an Apache's Open Source implementation of Google's MapReduce parallel processing framework. Running Hadoop on a cloud means that we have the facilities to add or remove computing power from the Hadoop cluster within minutes or even less by provisioning more machines or shutting down currently running ones. In this survey, we investigate some methods to improve scalability of Hadoop platform and Autoscaling of that. Based on the evaluation methods we understand that "The controller module and BEEMR" are best way to improve energy performance.

Read the paper · More papers on PaperTik