Random Forest-Based Load Prediction for Cloud Data Centers

Shuyue Yin · 2024

Because of the growing complexity of data structure and changing workload, resource management in data centers remains an important problem. Load forecasting is an essential prerequisite for resource management decisions in cloud data centers. In this paper, we propose an intelligent machine learning model for CPU utilization of cloud data centers, which aims to solve the problem of resource allocation. In this paper, a variety of machine learning methods are investigated, including Random Forest, Decision Tree, Support Vector Regression, Ridge Regression, K Nearest Neighbor, and Principal Component Analysis. This allows us to forecast CPU utilization by leveraging additional features. In addition, we discuss two types of thoughts on whether to use dimensionality reduction. Our findings suggest that our data has highly correlated features and non-linear structure, so using Principal Component Analysis for dimensionality reduction on the dataset is not effective. Among the integrated learning methods examined, Random Forest emerges as the top performer in predicting CPU utilization. In addition, it can estimate the importance of each feature.

Read the paper · More papers on PaperTik