Hadoop cluster monitoring and fault analysis in real time

Joey Pinto, Pooja Jain, Tapan Kumar · 2016

Failure of a task running on a Hadoop cluster is highly expensive in terms of computational time. A failure occurring even at the end phase of the task may cause the need to redo the entire task. Thus is really important to deploy fault tolerant techniques. Hadoop deploys a technique of checkpointing to prevent data loss. However, computational time-loss still pose a grim threat to critical applications. Hence a solution is proposed that uses SVM models trained with normal cluster resource usage statistics to predict and detect faults occurring in the cluster. The prediction engine can detect faults with minimal time delay and give sufficient time to implement fault tolerance and recovery measures.

Read the paper · More papers on PaperTik