A Survey of Fault Tolerance Techniques in Distributed Computing Environments

Dharmendra Kumar Tripathi, Manoj Kumar Niranjan, Yashawant Pathak, Pritaj Yadav · 2026

Distributed computing platforms underpin modern cloud services, edge computing deployments, and high-performance computing systems. As these systems scale and become increasingly heterogeneous, failures such as hardware faults, software defects, network partitions, and transient errors become unavoidable. Fault tolerance mechanisms therefore play a central role in sustaining availability and correctness. This paper survey&s;s fault tolerance techniques in distributed computing environments with a focus on checkpointing and rollback recovery. We classify checkpointing methods as uncoordinated, coordinated, and communication-induced, and relate them to log-based recovery. We also discuss how fault tolerance has evolved in cloud orchestration, containerized systems, and large-scale artificial intelligence and high-performance computing workloads. A comparative table summarizes trade-offs in coordination, overhead, storage, and recovery guarantees. Finally, we outline open research challenges including adaptive checkpointing, predictive failure handling, and resource-efficient resilience for next-generation distributed systems.

Read the paper · More papers on PaperTik