Horizontal Pod Autoscaling based on Kubernetes with Fast Response and Slow Shrinkage
Qizheng Huo, Shaonan Li, Yongqiang Xie, Zhongbo Li · 2022
Autoscaling is an important part of Kubernetes. Horizontal Pod Autoscaling (HPA) schedules cluster resources according to the service load status to ensure that services are still running normally when the load increases or decreases. But the default HPA is slow and inflexible. In scenarios with large load changes, it will scale at a high frequency, resulting in a waste of cluster resources and reducing the robustness of the cluster. This paper proposes a Fast-response and Slow-shrinkage algorithm. By setting the expansion rate, tolerance, window time, etc., the flexibility of scaling is increased. When facing the pressure of a sudden increase in traffic, the rapid expansion can be optimized in minutes. The algorithm also adopts a cautious shrinking strategy to reduce and prevent failures caused by secondary access peaks.