AI-Driven Predictive Auto-Scaling for Efficient Microservices Deployment in Cloud Data Centers using LSTM and Kubernetes HPA
M T Vasumathi, Manju Sadasivan, V Asha, Vema Venkata Nagasai Phanindra Kumar, Bhima Kalyan Kumar, V Veeresh · 2025
Cloud computing has revolutionized modern business operations, but efficient resource management remains a critical challenge. Traditional auto-scaling methods, such as Kubernetes Horizontal Pod Autoscaler (HPA), operate reactively, leading to resource inefficiencies, increased operational costs, and delayed scaling responses. This study proposes an AI-driven predictive auto-scaling framework utilizing Long Short-Term Memory (LSTM) networks to forecast workload demands and dynamically adjust resource allocation in Kubernetes-based micro services. The model anticipates scaling needs before workload surges occur, minimizing response delays and optimizing cloud resource utilization. Experimental evaluations demonstrate a 28% improvement in system response time and a 50% reduction in scaling delay compared to conventional HPA. Additionally, the proposed approach enhances CPU utilization efficiency and reduces energy consumption by 15%, promoting sustainability in cloud environments. These findings establish AI-driven predictive auto-scaling as a superior alternative to traditional methods, ensuring improved performance, cost-effectiveness, and environmentally sustainable cloud operations.