Reducing HPC energy footprint for large scale GPU accelerated workloads

Gabriel Hautreux, Etienne Malaboeuf · 2023

As the energy cost continues to rise, High Performance Computing (HPC) centers may seek to reduce their energy footprint. This could be implemented as a temporary production shutdown or in a more scientific production friendly way where the machine is able to increase its viability through increased efficiency. In this document we examine the second approach using a French, production ready machine hosted at Centre Informatique National de l'Enseignement Supérieur (CINES) in Montpellier. This machine, Adastra, is based on the AMD MI250X GPU architecture and currently #3 in Green500. Adastra is used by hundreds of French researchers, representing dozens of different applications from different scientific fields. As a base for the study, we define a set of applications representative of our current HPC and AI production workload. In this parametric study, we characterize our very diverse workload by applying a range of frequency capping or power capping policies at the node level in order to build an efficiency profile of each application. Based on the collected results, we produce guidelines trading between pure energy savings (energy to solution) to pure performance (time to solution) for each applications and, more importantly, for the production workload as a whole. We hope the results of this study will be of help to accelerators enabled HPC centers seeking to reduce their energy footprint by applying policies on either accelerators frequency or power capping at the node level.

Read the paper · More papers on PaperTik