Co-Locating Virtual Machine Logging and Replay for Recording System Failures
Jin Kawasaki, Shuichi Oikawa · 2010
There can be more system failures in the near future because of the combination of increased software complexity and a wide variety of usage patters. It is, however, difficult or sometimes almost impossible to find the root causes of the failures only from the limited and unreliable information provided by customers. Therefore, it is important to equip a feature that enables the complete tracing of system failures. We propose a system that employs two virtual machines, one for the primary execution and the other for the backup execution. The backup virtual machine maintains the past state of the primary virtual machine along with the log to make the backup the same state as the primary. When a system failure occurs on the primary virtual machine, the VMM saves the backup state and the log. By replaying the backup virtual machine from the saved state following the saved log, the execution path to the failure can be completely traced. We developed such a logging and replaying feature in a VMM. It can log and replay the execution of the Linux operating system. The experiments show that the overhead of the primary execution is only fractional, and the overhead of the replaying execution on the backup is less than 2%.