The Recovery System for Hadoop Cluster.
Priya Deshpande, Darshan Bora · Distributed Multimedia Systems · 2014
Due to brisk growth of data volume in many organizations, large-scale data processing became a demanding topic for industry as well as for academic fields. Hadoop is widely adopted in Cloud Computing environment for unstructured data. Hadoop is an open source, a java based distributed computing framework, and supports large-scale distributed data processing. In the recent years, Hadoop Distributed File System (HDFS) is popular for huge data sets and streams of operation on it. Availability of Hadoop is the important factor in Cloud Computing. But, in HDFS, Namenode failure affects the performance of the Hadoop cluster. It can be a single point failure. In this paper, we analysed the behaviour of Namenode, what are effects of Namenode failure. This paper presents a scenario to overcome this failure. Our scenario replicates the Namenode on the other Datanode so that the availability of the metadata is increases which will reduce the loss of data as well as delay. Keywords— Hadoop; Cloud Computing; HDFS; Namenode; availability; failure.