HADOOP SKELETON & FAULT TOLERANCE IN HADOOP CLUSTERS
Vishal S Patil, Pravin D. Soni · 2013
In the today’s era of information technology and computer science storing and processing a data is very important aspect. Nowadays even a terabytes and petabytes of data is not sufficient for storing large chunks of database. Hence companies today use concept called Hadoop in their application. Even sufficiently large amount of data warehouses are unable to satisfy needs of data storage. Hadoop is designed to store large amount of data sets reliably. Hadoop is popular open source software which supports parallel and distributed data processing. It is highly scalable compute platform. Hadoop enables users to store and process bulk amount which is not possible while using less scalable techniques. Along with reliability and scalability features Hadoop also provide faults tolerance mechanism by which system continues to function correctly even after some components fail’s working properly. Faults tolerance is mainly achieved using data duplication and making copies of same data sets in two or more data nodes. In this paper we describe the framework of Hadoop along with how fault tolerance is achieved by means of data duplication.