An In-Memory Checkpoint-Restart Mechanism for a Cluster of Virtual Machines

Jumpol Yaothanee, Kasidit Chanchio · 2019

A cluster of virtual machines can be used to execute parallel applications in Cloud Computing environments. However, the cloud infrastructure may fail at any time for a variety of reasons. Although a coordinated checkpointing capability at the hypervisor level is highly transparent to parallel applications, existing solutions still suffer from excessive checkpoint time and downtime. They also cause significant application execution delays due to packet loss. This paper introduces IMVCCR, a novel in-memory Checkpoint-Restart mechanism for a virtual cluster. IMVCCR consists of a framework that performs coordinated checkpointing for the entire cluster. It reduces checkpoint time and downtime by applying live migration and using main memory as transient checkpoint storage. IMVCCR also uses an efficient synchronization mechanism to reduce packet loss. Preliminary experiments show that IMVCCR generates very low checkpoint times and downtimes. It also incurs low overheads in the total execution time of parallel applications.

Read the paper · More papers on PaperTik