Evaluation of Kubernetes Schedulers for a Community Cloud Computing Model
Erik S. Gough · 2024
Providing access to computational resources for scientific computing is a cornerstone of research institutions. In turn, shared computing models have been developed to provide researchers access to cost effective and reliable computing resources with less overhead than research lab scale systems. These models have traditionally been centered around high performance computing (HPC) architectures, using batch systems for processing jobs with well defined resource requirements and run times. Purdue’s collaborative computing model, the Community Cluster Program, has served Purdue researchers, faculty and staff for nearly two decades and is the the reference "condo" HPC model followed by many other universities. However, the architectures provided by the Community Cluster Program do not meet the requirements of new data analysis methods and applications that were born in the cloud. In recent years, we have seen the rise Kubernetes as a platform for scientific computing with many universities now hosting their own Kubernetes resources alongside HPC infrastructure. While Kubernetes has grown in adoption, there are still limiting factors in current deployment methods that prevent a shared condo-like environment and effective resource sharing. In this paper, we evaluate multiple Kubernetes schedulers to determine their resource sharing effectiveness as part of a "Community Cloud" computing model. Results show Kubernetes schedulers like the capacity scheduling plugin and YuniKorn can reduce workflow run times up to 4.6x and provide a 3x increase in overall cluster utilization.