HeteroScheduler: Dynamic Task Scheduling for CPU-GPU Optimization and Contention Mitigation in Cloud Data Centers
Seokwon Choi, Hyeonsang Eom · 2025
Cloud data centers are increasingly integrating high-performance servers, shifting to heterogeneous environments where maintaining homogeneous infrastructures becomes inefficient. Thus, GPU-CPU heterogeneous cloud data centers have become essential. Optimal scheduling is essential to handle server performance differences and resource usage variations. Efficient resource utilization in heterogeneous clusters is a critical challenge for optimizing task execution performance and overall system efficiency. Traditional scheduling techniques primarily rely on offline static profiling, measuring task execution time in advance for each device, or classifying tasks based on CPU and memory utilization to colocate workloads with different characteristics, thereby mitigating resource contention and improving performance. However, these approaches fail to dynamically respond to fine-grained resource contention and environmental changes, leading to performance degradation and reduced resource efficiency. To address this issue, this study proposes a dynamic scheduling framework that utilizes real-time resource metrics. The proposed framework classifies tasks into Compute-intensive and Memory-intensive categories to determine optimal task placement in a heterogeneous cluster environment. Compute-intensive tasks require high IPC (Instructions Per Cycle) and primarily utilize CPU resources, while Memory-intensive tasks exhibit high LLC miss rates and consume substantial memory bandwidth. After defining task classification metrics through extensive experiments, the framework assigns tasks to either CPU or GPU servers based on their characteristics. Additionally, real-time resource monitoring detects contention and triggers task migration to optimize resource utilization, mitigate contention, and maximize performance. Experimental results demonstrate that the proposed method outperforms existing heterogeneous cluster scheduling techniques, achieving notable improvements in resource utilization, task execution time, and contention mitigation.