Evolutionary Computation-Based Scheduling of Machine Learning Workloads for GPU Clusters
Seokmin Kwon, Hyokyung Bahn · 2024
Recently, machine learning (ML) workloads across diverse industries such as smart logistics, finance, and entertainment are increasingly being executed on cloud platforms. Efficient scheduling of these ML workloads is challenging as various types of workloads coexist and the cluster systems feature heterogeneous GPU resources. Although task scheduling has been extensively studied, traditional scheduling policies do not perform well in such environments as they cause resource fragmentation problems, which significantly lowers GPU utilization. To address this issue, this paper proposes a new scheduling approach utilizing evolutionary computation techniques, and implements it within a process-based event simulation framework. Experimental results, replicating extensive ML task traces collected from Alibaba’s MLaaS cluster, demonstrate that the proposed scheduling approach significantly improves GPU utilization compared to conventional scheduling policies. It is anticipated that the scheduling policies proposed in this paper will be used effectively for the resource allocation of ML workloads in future GPU cluster systems.