Shiraz: Exploiting System Reliability and Application Resilience Characteristics to Improve Large Scale System Throughput
Rohan Garg, Tirthak Patel, Gene D. Cooperman, Devesh Tiwari · 2018
Large-scale applications rely on resilience mechanisms such as checkpoint-restart to make forward progress in the presence of failures. Unfortunately, this incurs huge I/O overhead and impedes productivity. To mitigate this challenge, this paper introduces a new technique, Shiraz, which demonstrates how to exploit differences in the checkpointing overhead among applications and knowledge of temporal characteristics of failures to improve both the overall system throughput and performance of individual applications.