Machine Learning Model Update Strategies for Hard Disk Drive Failure Prediction
Marwin Züfle, Florian Erhard, Samuel Kounev · 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) · 2021
The growing size of today’s data centers and the expectation of 24/7 availability continuously increase the complexity of hardware administration. To this end, the Self-Monitoring, Analysis, and Reporting Technology has been developed to provide insights into the health state of hard disk drives. Many approaches to predicting hard disk drive failures based on such monitoring data have been proposed in recent years. Nevertheless, most approaches consider this problem only as a static task, i.e., they train a static machine learning model on a given training set and evaluate its performance on a test set. However, due to model aging and changes in failure patterns, previously learned prediction models must be updated during runtime, which requires a time-dependent evaluation. Therefore, we present four machine learning model updating strategies, build multiple models for hard disk drive failure prediction using four machine learning algorithms, and compare the prediction quality of the different model update strategies and machine learning algorithms. Experimental results using a real-world data set of hard disk drives demonstrate the need for model update strategies, with XGBoost using the Hoeffding bound update trigger achieving the overall best prediction performance concerning prediction quality and number of updates required.