A deep learning approach for early prediction of task failures in cloud computing environments
Saba Aldomi, Husam Suleiman, Ali M. Shatnawi, Luay M. Alawneh · Systems and Soft Computing · 2026
Prediction of task failures in cloud computing is of great importance due to its critical impact on task execution and resource utilization. Potential risks associated with failure events of tasks can lead to dissatisfaction among clients relying on cloud services. Therefore, it is crucial to comprehend properties and attributes of task failures to prevent them, or at least, to develop the capability to tolerate them. While there has been research conducted on failure analysis, there is a notable lack of emphasis on the application of Artificial Intelligence (AI) in characterizing and predicting task failures. This study aims to address this gap by developing a failure prediction framework capable of early identification of failed tasks and predicting the type of failure events that represent task states throughout their life cycle. We present a hybrid feature extraction and classification framework that uses SelectKBest for feature pre-selection and a GRU network as a sequence-level feature extractor. The extracted features are utilized by the GRU to train machine learning classifiers for the multi-class prediction phase including Random Forest (RF), K-Nearest Neighbor (KNN), and Support Vector Machine (SVM). The framework presents several benefits, including reducing resource wastage and Service Level Agreement (SLA) violations. The framework is evaluated based on the analysis of Google cluster traces in which the task states are Enable, Evict, Lost, Finish, Kill, Fail, Queue, Schedule, Update Pending, and Update Running. The findings show that a GRU model trained with the top 14 features achieves a test accuracy of 97.7% for feature extraction and that the combined GRU-RF yields the best predictive performance (overall RMSE = 0 . 1415 , Fail-class F1 = 0 . 99 % , average AUC per class > 0 . 98 ).