A Cloud-Native Heterogeneous Resource Scheduling Platform Supporting GPU Scheduling
Hongjie Liu, X.Z. Wang, C. Wang · 2025
This paper proposes and implements a cloud-native heterogeneous resource scheduling platform based on Kubernetes and Docker, aimed at optimizing the scheduling and utilization of GPU resources in cloud environments. The platform integrates Kubernetes' container orchestration capabilities, NVIDIA GPU device plugins, and a custom scheduling algorithm to achieve dynamic management of GPU resources, effectively balancing the load of GPU-intensive tasks. Furthermore, the platform supports hybrid scheduling of GPUs and CPUs, optimizing task allocation among heterogeneous resources and enhancing overall computational efficiency. The system incorporates Prometheus and Loki for real-time monitoring and log management, assisting administrators in maintaining resource usage, detecting bottlenecks, and ensuring system stability. The platform leverages Kubernetes to provide self-healing capabilities, ensuring high availability in the event of node or container failures. Through extensive testing and validation, the proposed platform demonstrates outstanding performance in terms of functionality, efficiency, and fault tolerance, making it suitable for application scenarios involving GPU-intensive tasks such as deep learning, scientific computing, and image rendering.