Transparent Fault Tolerance for Stateful Applications in Kubernetes with Checkpoint/Restore
Henri Schmidt, Zeineb Rejiba, Raphael Eidenbenz, Klaus-Tycho Förster · 2023
This paper presents a solution providing fault tolerance for stateful containerized applications that is transparent, i.e., the application does not require to structure or manage its state in any particular fashion. In the case of faults, such as node crashes or node isolation, the application resumes execution on another node. The solution relies on a Kubernetes operator and a tool to periodically checkpoint containers and restore from the latest checkpoints in case of a node failure. Experimental evaluations reveal the trade-offs between over-head due to checkpointing, i.e., CPU load, memory, network bandwidth, reduced availability, and the performance during recovery, i.e., outage time, state quality. Compared to a non-transparent solution, the transparent solution yields similar downtimes and state quality at an increased overhead.