Practical Difficulties and Anomalies in Running Hadoop
Dick Mugerwa, Jeong Geun Ji, Youngmi Kwon · 2017
Currently, Hadoop is dominantly used framework for processing large data. It has also been noted that Hadoop is one of the most popular implementation of MapReduce which is predominantly used by Facebook and Yahoo for large data parallel distributed computing. A typical Hadoop cluster comprises of single name node and many data nodes, with an assumption that all the nodes in a cluster are homogeneous meaning that all nodes are expected to finish computing a job at the same time. But in a heterogeneous cluster, data nodes have a varying range of computation capabilities. Many researches have come up with different algorithms to overcome the challenges in Hadoop to improve service time of jobs and resource utilization. This paper addresses the practical difficulties and anomalies in running Hadoop system built in LINUX CENT OS 6.6. Up to six data nodes are configured in our Hadoop architecture. Different RAM sizes showed unexpected result and MapReduce process showed some anomalies which have not addressed in documentation.