Optimizing VM Provisioning of MapReduce Tasks on Public Cloud
Banpreet Kaur, Ankit Grover · 2016
Big data has been one of the important reaseach topics in recent years. People from industries are realizing the value of explosively growing datasets. These huge datasets require the use of parallel processing for faster computation. MapReduce framework has received a wide acclaim over the past few years for large scale computing of Big Data. Map reduce divides a job into map and reduce tasks and performs processing on them parallely Cloud Computing is emerging as a new paradigm for computation. It provides elastic computing and Storage resource on demand. Cloud Computing resources are best fit for big data processing as they allow applications to scale at runtime. Processing big data on cloud involves optimizing the provisioning of resources so as to meet SLA requirement of Clients. The paper addresses this issue by providing a cost effective resource provisioning approach and at the same time increases the overall utilization of physical machines by using the machine to its maximum extent and putting the low utilized physical machines to standby mode, by migrating the load on to other physical machines.