A LAHC-based Job Scheduling Strategy to Improve Big Data Processing in Geo-distributed Contexts
Marco Cavallo, Giuseppe Modica, Carmelo Polito, Orazio Tomarchio · 2017
The wide spread adoption of IoT technologies has resulted in generation of huge amount of data, or Big Data, which has to be collected, stored and processed through new techniques to produce value in the best possible way.Distributed computing frameworks such as Hadoop, based on the MapReduce paradigm, have been used to process such amounts of data by exploiting the computing power of many cluster nodes.Unfortunately, in many real big data applications the data to be processed reside in various computationally heterogeneous data centers distributed in different locations.In this context the Hadoop performance collapses dramatically.To face this issue, we developed a Hierarchical Hadoop Framework (H2F) capable of scheduling and distributing tasks among geographically distant clusters in a way that minimizes the overall jobs execution time.In this work the focus is put on the definition of a job scheduling system based on a one-point iterative search algorithm that increases the framework scalability while guaranteeing good job performance.