Optimizing VM provisioning of map reduce tasks on public clouds
Banpreet Kaur, Nidhi Jain · 2016
Today's fast data generation from massive sources is calling for efficient big data processing, which imposes huge demands on the networking Infrastructures and computing. MapReduce suits very well to develop applications that can process large amount of data in a distributed fashion over large clusters. While there have been studies on improving performance of MapReduce in dedicated server clusters, the research in the context of public cloud is still in progress. One of the main issue of running map reduce tasks on public clouds is optimize resource provisioning to minimize the cost or job finish time for a specific job. The paper addresses this issue by providing a cost effective resource provisioning approach and at the same time increases the overall utilization of physical machines by using the machine to its maximum extent and putting the low utilized physical machines to standby mode, by migrating the load on to other physical machines.