Enhancing Fault Tolerance System in Cloud Environment Through Failure Prediction Using Ensemble Learning
K. Vani, S. Sujatha · 2024
Cloud environments are dynamic, with changing workloads, configurations, and system states. Adapting to these dynamics poses a challenge for fault tolerance systems. Imbalances in the distribution of failure and non-failure instances can lead to biased models that struggle to accurately predict less frequent failure events. Most cloud services, both logical and physical, have failed as a result of the size and variety of the cloud. The behaviors of successful and failed actions using the traces that are now available are categorized and evaluated. A strategy to predict which jobs will fail was developed and put into practice. The prior research applied the traditional machine learning algorithm for job failure prediction. However the existing machine learning fails to capture the characteristics of the jobs and produces the result with less accuracy and also with high error rate. The ensemble machine learning algorithms is utilized for job failure forecasting on Google cluster dataset. The suggested approach more effectively uses resources while optimizing cloud applications. This technology can be used by large data centers to identify unsuccessful jobs before the cloud management system schedules them. The present model is compared with prior studies with respect to accuracy and error rates. The results exposed that the proposed prediction model has high precision, recall, and Fl-score.