Privacy preserving scheme for MapReduce

Madhvaraj M. Shetty, D. H. Manjaiah · 2017

In present world, each and every industry/organization is generating a large volume of data. Capturing, processing and analyzing this vast amount of data is very difficult using traditional database and other tools. Meantime, the transformation of cloud services presents great opportunity for organizations to increase productivity at the same time reducing their cost. Cloud computing offers variety of services for processing data on multiple large-scale datasets with the help of distributed computation framework. One of the most common platform-as-a-service in cloud computing is MapReduce, which is introduced by Google in 2004. It is a programming model that provides and efficient parallel and fault tolerant processing on large volume of data without requiring any costly dedicated computing nodes. It allows to process huge amount of data by distributing data as well as computing over a number of nodes. MapReduce is the key approach to answer today's Big Data problems because of its scalability. Today it is used widely around the world as an efficient distributed computation tool to solve large class of problems. However, due to scattered computation and distributed sensitive data among various nodes and the flow of sensitive data between the nodes while computation, the security and privacy concerns in MapReduce are aggravated. In the last decade, the privacy protection and security of data is the one of the most concerned issue in the cloud applications and in big data platform. In this paper, we discussed about MapReduce with its privacy and security aspects. Then we develop a solution which maintains confidentiality intermediate data as well as the data stored in HDFS which can preserve data privacy.

Read the paper · More papers on PaperTik