Performance-Reliability Modeling for Long-Running Programs
Marina Khludova · 2018
This paper studies the execution of a long-running program in an unreliable system that can eventually fail, and then it is being repaired, but all performed work is lost as the system fails and the computation has to start anew. Practitioners are now interested in a redundancy mechanism which should introduce small overhead and is aimed at reducing the amount of work that is lost upon failure of the system. In connection with such interest, bounded accumulated unproductive time is considered and an expression is obtained for the moment-generating function of this random variable in the case of a successful program completion. Rigorous mathematical models for detailed studies of the accumulated unproductive time under different operating conditions are presented. Conducted studies indicate that the mechanism of checkpointing can significantly reduce such metric as average loss factor (based on unproductive time), but in some cases it introduces an unacceptable overhead depending on the strategy used for its implementation.