MapReduce scheduling in hybrid cloud with multi-level privacy
Toon Degryse, Sucha Smanchat · 2015
MapReduce is a popular programming model for analyzing large datasets. It allows for parallel processing of large amounts of data over commodity computers or cloud computing resources. Hadoop is an open-source implementation of MapReduce and uses the Capacity or Fair scheduler to assign map and reduce tasks to the various computing nodes. They both address fair share of resources between the users but do not consider any privacy constraint, an important concern when using cloud resources due to clouds generally having a different risk of exposure. Jobs are currently sent to cloud resources disregarding the different privacy implementations of each cloud involved. This paper suggest a multi-level privacy scheduler with data clearance levels and cloud authorization levels so the mapping of map and reduce tasks to cloud resources can satisfy the organization's privacy constraints.