Power-Efficiency Variation on A64FX Supercomputers and its Application to System Operation

Tomoya Kusaba, Yusuke Awaki, Kohei Yoshida, Shinobu Miwa, Hayato Yamaki, Toshihiro Hanawa, Hiroki Honda · 2024

Understanding the variation in power-efficiency between compute nodes is important for energy-efficient operation in modern supercomputing systems. A64FX supercomputers have various distinctive features in power management and are therefore considered to exhibit the unique nature of the power-efficiency variation, which may be useful for energy-efficient system operation. However, previous studies have not discussed the difference in power-efficiency variation between applications running on compute nodes. In this paper, we first analyze the impact of applications on the variation in power-efficiency on A64FX supercomputers. Through the analysis of our experimental data collected by executing various applications on 12,289 nodes in Fugaku and 6,144 nodes in Wisteria/BDEC-01 Odyssey (referred to as Wisteria-O), we found that the order of the power efficiency of the compute nodes was almost independent of the types of applications running. To the best of our knowledge, this feature is very unique to A64FX supercomputers and has not been observed in other systems. Based on this observation, we propose a variation-aware method that reduces the number of operating nodes to maximize the compute capability of a supercomputing system under a given power constraint. Our approach ranks all compute nodes with their power efficiency and classifies nodes with similar power efficiency into a group in advance by physically relocating them. When we reduce the number of operating nodes, we repeatedly select the most power-inefficient node group as shutdown nodes until the system's power consumption meets a given constraint. Our experimental results show that the proposed approach can increase the compute capability in Fugaku and Wisteria-O by up to 13% and 21% under power constraints of 0.22 MW and 0.11 MW, respectively, compared to the worst-case scenario.

Read the paper · More papers on PaperTik