A Survey on Failure Prediction in Large-scale Computing Systems
Fei Xia, Hu Song, Long Yan, Yan Li, Lijun Wang · 2021
With the development of many data-intensive applications, large-scale systems have been widely used to solve advanced computation problems. As the increasing growth of complexity and scale, these systems are more likely to confront failure. Since remedial measures for failure take huge cost and effort at today's large-scale systems, fault tolerance which aims to decrease the impacts of fault on systems has become a necessity instead of an option. As one of the key techniques of fault tolerance, failure prediction has made itself an increasing important issue to improve the resource efficiency and the availability of systems. Over the past few years, a multitude of innovative failure prediction approaches has emerged such as mathematical and statistical modeling, machine learning techniques and so forth. Unfortunately, they are currently so poorly classified that it is difficult to figure out the wide spectrum of methods concerning with this area. To this end, we provide an extensive and comprehensive survey of existing research work in the area of failure prediction via exploring and analyzing over 20 various approaches. Also, we develop our own taxonomy assisting in classifying methods, which makes it easier to understand and compare the pros and cons of these methods in respective category.