Cloud Platform Optimization Using Azure Machine Learning to Improve Performance and Reliability
S Senthil Pandi, Prashant Kumar, T A Salman Latheef · 2024
Cloud platforms are crucial components of contemporary computer infrastructures, necessitating strong availability and resilience to guarantee continuous service delivery. Unfortunately, current systems frequently have trouble resolving problems on their own in real time, which causes downtime and performance deterioration. In order to improve cloud platform resilience, we suggest an integrated strategy that makes use of machine learning (ML) inside the Azure ecosystem. Our process includes gathering metadata from computing units and load balancers, developing machine learning models to forecast traffic patterns, and putting these models into use to help anticipate problems and make wise decisions. Experimental simulations used for evaluation show notable gains: the ML-enabled architecture achieves recovery times that are quicker on average and reduces error rates by during failures when compared to the current systems. The approach described in this paper is a multi-step procedure that begins with the extraction of metadata from necessary cloud platform components, such load balancers and compute units. Based on traffic estimates and performance indicators, these metadata are used to train machine learning (ML) models, which are subsequently used to make informed judgments and take proactive remedial action. The incorporation of machine learning techniques into cloud platforms is a revolutionary strategy for tackling the problems of system resilience and downtime. Our suggested design provides a scalable and effective way to proactively identify and mitigate any problems by utilizing Azure’s infrastructure and ML capabilities. This will eventually ensure continuous service delivery and improve user experience.