Deep Learning-based fault prediction in cloud system
Dinh-Dai Vu, Xuan Tuong Vu, Younghan Kim · 2021 International Conference on Information and Communication Technology Convergence (ICTC) · 2021
A cloud system typically contains a huge number of computing nodes with numerous services running on. Nodes can fail in practice, leading to the downtime of services, hence, the reliability of cloud service is a crucial issue for cloud service providers. To make the cloud environment more reliable, predicting the failure before it already happens is important. The ability to predict faulty nodes enables the migration of service to the healthy nodes, therefore improving service availability. Proactive fault prediction techniques with deep learning methods that are based on historical metrics (such as CPU usage, memory usage, etc) to predict future failures can be used to solve effectively this problem. In this paper, we propose to use the Bidirectional Long Short Term Memory (Bi-LSTM) model for predicting the failure-proneness of nodes in a cluster based on time series data.