An Auto Scaling Energy Efficient Approach in Apache Hadoop
Nemouchi Warda Ismahene, Souheila Boudouda, Zarour Nacereddine · 2020
Cloud Computing has emerged as revolutionary paradigm for large-scale data intensive analysis over the last decade. In addition, Map Reduce and its implementation Hadoop have been successful at developing and running Big Data Distributed computations. However, their effect on datacenters energy efficiency has become significant; some of the servers are run without being used actively on daily basis. Making use of Cloud Computing advantages such as elasticity and scalability along with Hadoop's powerful distributed architecture has been an important research axis. The ability of managing resources (adding/removing nodes that run Map Reduce jobs to the cluster) automatically based on workloads without affecting time response has been investigated. This paper presents an approach of auto-scaling in the Hadoop framework, we have focused on separating nodes to core/computation to avoid data loss and guarantee the ability to remove nodes smoothly and instantly.