OME: An Optimized Modeling Engine for Disk Failure Prediction in Heterogeneous Datacenter

Yanwen Xie, Dan Feng, Fang Wang, Xinyan Zhang, Jizhong Han, Xuehai Tang · 2018

Nowadays, there are lots of disks from various disk models in datacenter. It is a challenge to make failure prediction for all disk models with high precision and high coverage. One-for-one modeling, transfer learning modeling and one-for-all modeling are proposed to address the challenge. However, none of them works well for all disk models and the automation problem for method selection and parameter tuning still persists. In this paper, we propose OME, an optimized modeling engine for disk failure prediction in heterogeneous datacenter. It builds a basis predictive model with one-for-all modeling and searches for the optimized with one-for-one and transfer learning modeling for every disk model. To achieve automation, OME employs a simple but effective transfer learning method, does cross-validation for comparison, prunes the tuning space, and constructs a directed acyclic graph for parallelism. Evaluation on a dataset from a real-world datacenter shows that OME outperforms a one-for-all predictive model from previous work by 18.5% overall, and the improvement for 43.3% disk models reaches over 30%.

Read the paper · More papers on PaperTik