Efficient Deterministic Replay through Dynamic Binary Translation
Piyus Kedia · 2015
We present an efficient software implementation to deterministically record and replay a full multiprocessor virtual machine (VM), including its guest OS kernel and applications. Deterministically replaying a shared-memory monolithic OS kernel (like Linux) presents a significant performance challenge, and we demonstrate the use of dynamic binary translation to achieve this objective. Dynamic binary translation (DBT) is a powerful technique with several important applications. System-level binary translators have been used for implementing a Virtual Machine Monitor [2] and for instrumentation in the OS kernel [29]. In current designs, the performance overhead of binary translation on kernel-intensive workloads is high. e.g., over 10x slowdowns were reported on the syscall nanobenchmark in [2], 2-5x slowdowns were reported on lmben h microbenchmarks in [29]. These overheads are primarily due to the extra work required to correctly handle kernel mechanisms like interrupts, exceptions, and physical CPU concurrency. Since the overhead of DBT is itself very high hence we can not use it for improving determinstic replay performance. We present a kernel-level binary translation mechanism which exhibits near-native performance even on applications with large kernel activity. Our translator relaxes transparency requirements and aggressively takes advantage of kernel invariants to eliminate sources of slowdown. We have implemented our translator as a loadable module in unmodified Linux, and present performance and scalability experiments on multiprocessor hardware. Although our implementation is Linux specific, our mechanisms are quite general; we only take advantage of typical kernel design patterns, not Linux-specific features. The biggest challenge in deterministically replaying a multiprocessor system is recording the order of shared memory read and writes. The previous comparable approach [27] uses CREW (Concurrent Read Exclusive Write) protocol at the page granularity. Page grained CREW protocol uses hardware page protection technique(EPT/Shadow) to restrict the access privilege of the CPUs such that, multiple CPUs can read from a page by acquiring shared access privilege of the page but for writing to a page they needs to acquire the exclusive access privilege of that page. This scheme suffers from false sharing and huge shuttling between processors for benchmarks having large amount of sharing such as Linux kernel. Every transfer