Cost-Performance of Fault Tolerance in Cloud Computing
Yong Meng Teo, B.L. Luong, Yingwen Song, Tang Quoc Nam · 2011
As more computation moves into the highly dynamic and distributed cloud, applications are becoming more vulnerable to diverse failures. This paper presents a unified analytical model to study the cost-performance tradeoffs of fault tolerance in cloud applications. We compare four main checkpoint and recovery techniques, namely, coordinated checkpointing, and uncoordinated checkpointing such as pessimistic sender-based message logging, pessimistic receiverbased message logging (PR), and optimistic receiver-based message logging (OR). We focus on how application size, checkpointing frequency, network latency, and mean time between failures influence the cost of fault tolerance, expressed in terms of percentage increase in application execution time. We further study the cost of fault tolerance in cloud applications with high probability of failure, network latency and message communication among executing processes (virtual machines). Our analysis shows that the cost of fault tolerance for both OR and PR is around 5 % for the range of application size (32 to 4,096 processes or virtual machines), checkpointing frequency (one checkpoint per minute to one checkpoint per hour) and network latency (Myrinet to Internet) evaluated. For a cloud with high network latency, the cost of fault tolerance is about 5 % in both OR and PR, but when failure probability is low, OR is a suitable choice and when failures are more frequent, PR is a better candidate. 1