Exploring Non-Volatility of Non-Volatile Memory for High Performance Computing Under Failures
Jie Ren, Kai Wu, Dong Li · 2020
Hardware failures and faults often result in application crash in HPC. The emergence of non-volatile memory (NVM) provides a solution to address this problem. Leveraging the nonvolatility of NVM, one can build in-memory checkpoints or enable crash-consistent data objects. However, these solutions cause large memory consumption, extra writes to NVM, or disruptive changes to applications. We introduces a fundamentally new methodology to handle HPC under failures based on NVM. In particular, we attempt to use remaining data objects in NVM (possibly stale ones because of losing data updates in caches) to restart crashed applications. To address the challenge of possibly unsuccessful recomputation after the application restarts, we introduce a framework EasyCrash that uses a systematic approach to automatically decide how to selectively persist application data objects to significantly increase possibility of successful recomputation. EasyCrash enables up to 30% improvement (20% on average) in system efficiency at various system scales.