An approach for log analysis based failure monitoring in Hadoop cluster

Madhury Mohandas, P. M. Dhanya · 2013

Massive and gargantuan amount of data is produced on per day basis. Such scenario elevates the need for apposite storage, supervision and processing of data. The massive use of Distributed framework calls for faster analysis and diagnosis of failures. Due to the distributed nature of processing, it is difficult for cluster administrator to isolate the failures and failed nodes. Many contributions have been done for failure monitoring, analysis etc in the last few years. Apache Hadoop's Jobtracker, Namenode, Secondary Namenode, Datanode and Tasktracker all generate logs. This paper aims at building a failure monitoring system from the scratch, by parsing and analyzing the Hadoop log files generated in the cluster. The monitoring system gives all relevant details related to the application, and points out the specific reason for failure, that is, whether an application failure or a network failure (these are the most common failures in the cluster).

Read the paper · More papers on PaperTik