Distributed Data-Parallel Programs from Sequential Data Processing in Cloud
S. Divya, Kevin Stella · SSRN Electronic Journal · 2016
Map-Reduce are a programming model that enables easy development of scalable parallel applications to process vast amounts of data on large clusters of commodity machines. Based on this new framework, we perform extended evaluations of Map Reduce-inspired processing jobs on an IaaS cloud system and compare the results to the popular data processing framework Hadoop. Through a simple interface with two functions, map and reduce, this model facilitates parallel implementation of many realworld tasks such as data processing for search engines and machine learning. In recent years ad hoc parallel data processing has emerged to be one of the killer applications for Infrastructure-as-a- Service (IaaS) clouds. Major Cloud computing companies have started to integrate frameworks for parallel data processing in their product portfolio, making it easy for customers to access these services and to deploy their programs. Nephele is the first data processing framework to explicitly exploit the dynamic resource allocation offered by today’s IaaS clouds for both, task scheduling and execution. Particular tasks of a processing job can be assigned to different types of virtual machines which are automatically instantiated.