On the Need for Training Failure Prediction Algorithms in Evolving Software Systems

Ivano Irrera, Joao A. Duraes, Marco Paulo Amorim Vieira · 2014

Failure prediction is a promising technique to improve dependability of computer systems, in particular when it is important to foresee incoming failures and take corrective actions to avoid downtime or data corruption. Failure prediction is especially adequate in long running systems where internal errors accumulate and eventually lead to failures. The problem is that such systems do evolve. The workload and even the system itself changes over time, and this may affect the performance of the failure predictor. However, training failure prediction algorithms is a complex and time-consuming task and should be performed only when needed. Thus, it is important to understand if a system change affects prediction performance, to avoid running the target system with an ineffective predictor and prevent unnecessary retraining efforts. In this work we study the performance of a failure predictor when used to forecast failures in a web-serving system subject to successive updates. We observe and analyze the variation of performance in terms of ROC-AUC using fault injection and virtualization for the generation of the data needed for the assessment. Our results suggest that re-training is indeed necessary.

Read the paper · More papers on PaperTik