Checkpointing Period Optimization of Distributed Fail-Operational Automotive Applications
Philipp Weiß, Emil Daporta, Andreas Weichslgartner, Sebastian Steinhorst · 2021
Achieving a cost-efficient fail-operational behavior of safety-critical software is crucial for autonomous systems. However, most applications hold a state such that a checkpoint is required to enable a safe recovery. Here, the challenge is to find the maximum possible checkpointing period while minimizing network and computational overhead. For this purpose, we present an approach to analytically derive the maximum checkpointing period by giving an upper bound on the number of missed computational steps due to failure effects. Worst-case results of our case study using a SLAM application are consistent with our analytically derived exact bound. Overall, by using our approach, a maximum achievable checkpointing period can be determined to reduce network overhead in order to achieve a cost-efficient and safe behavior of autonomous systems.