A Design and Implementation of Cluster Heartbeat Network for Efficient Fault Detection
Ahmad Shukri Mohd Noor, Emma Ahmad Sirajudin · Journal of Telecommunication Electronic and Computer Engineering (JTEC) · 2016
To achieve fault tolerance in a server cluster, fault detection capability is a primary prerequisite. Efficient fault detection is prompt, correct and complete. This paper revisited the technique called Reactive Failure Detection (RFD) that dynamically predicts a heartbeat delay from a cluster node. We also identified the requirements to deploy RFD in actual servers. A new cluster heartbeat network with concurrency is proposed to use push and pull interaction during live monitoring and determining node’s status. The prototype of the new model is tested on a platform running multiple independent web applications and analyzed for its implementation and design correctness.