Anatomy of Hadoop Mapreduce Execution

Parvathy Gopakumar, Neethu Maria John · IOSR Journal of Computer Engineering · 2017

Storing and monitoring Big data in widely distributed environments for 24/7 is a huge task for global service organizations.These datasets require high processing power which can't be offered by traditional databases as they are stored in an unstructured format.Apache Hadoop is open source software for reliable, scalable and distributed computing.This framework is inspired by Google's MapReduce structure in which application is broken down into numerous small parts and each part can be run in any node in the cluster.This paper contains detailed study of the execution of MapReduce programs over Hadoop cluster .It also discusses how the Hadoop platform offers an easy way of distributed Bigdata computing.

Read the paper · More papers on PaperTik