Introducing an Intelligent MapReduce Framework for Distributed Data Processing in Clouds

Amir H. Basirat, Asad I. Khan · 2013

The world of big data is in need of high levels of scalability and the question, how to effectively process large-scale data sets is becoming increasingly relevant. Furthermore the existing data management schemes do not work well when data is partitioned among numerous available nodes dynamically. Approaches towards parallel data processing in cloud, which offer greater portability, manageability and compatibility of applications and data, are yet to be fully-explored. With this in mind, in this paper we would like to explore the possibility to evolve a new type of data processing approach that will efficiently partition and distribute data for clouds. For this matter, loosely-coupled associative techniques, not considered so far, can be the key to effectively partitioning and distributing data in the clouds. Ability to partition data optimally and automatically will allow elastic scaling of system resources and remove one of the main obstacles in provisioning data centric software-as-a-service (SaaS).

Read the paper · More papers on PaperTik