Automatic Energy-Efficient Job Scheduling in HPC: A Novel SLURM Plugin Approach
Anders Aaen Springborg, Michele Albano, Samuel Xavier‐de‐Souza · 2023
This paper presents a novel approach to enable energy-efficient job scheduling in High-Performance Computing (HPC) environments through application-specific energy models. We propose an architecture that decouples scheduling heuristics to a Python plugin of the HPC scheduler SLURM. The approach leverages the principles of Service-Oriented Architecture and Clean Architecture to create a proof-of-concept system that is adaptable for production setups, providing a platform for integrating various energy-efficient scheduling models. We demonstrate the approach in a single-node HPC system with an energy saving of 11% for the High-Performance Conjugate Gradients (HPCG) benchmark, which represents modern applications’ data access patterns and computation. The proposed approach opens up possibilities for more complex setups, such as automatically scheduling jobs when energy is cheap and renewable, a practice already used in companies utilizing HPC.