AI and HPC Applications on Leadership Computing Platforms: Performance and Scalability Studies

JaeHyuk Kwack, Colleen Bertoni, Umesh Unnikrishnan, Riccardo Balin, Khalid Hossain, Yasaman Ghadar, Timothy J. Williams, Abhishek Bagusetty, Mathialakan Thavappiragasam, Väinö Hatanpää, Archit Kumar Vasan, John Robert Tramm, Scott Parker · 2025

As HPC systems move into the exascale era an increasing diversity of processing hardware is being deployed. The last decade saw the ascendance of NVIDIA GPU-accelerated systems among the largest scale HPC systems and spurred the need for application developers to consider approaches to performance portability that preserved developer productivity. This challenge has been compounded in the last several years by the introduction of the first two exascale systems, Frontier and Aurora (\#2 and \#3 on the November 2024 Top 500 list respectively). These systems utilize new and different GPUs, with the AMD MI-250X GPU on Frontier and the Intel Data Center GPU Max 1550 on Aurora. This study investigates the performance and qualitative performance portability of$\mathbf{1 2}$HPC and ML applications on three large scale HPC systems that utilize GPUs from the three different vendors: Frontier (AMD), Aurora (Intel), and Polaris (NVIDIA A100). The performance of these applications is evaluated on single GPU, single node, and multinode scales on each of the systems. We show that the figures-of-merit (FOMs) of the applications on a single GPU of Aurora and Frontier ranged from$0.9-4 x$and$0.8-2.5 x$, respectively, the performance on a GPU of Polaris. We also show that the FOMs on a single node of Aurora and Frontier ranged from 1.3-6.3x and 0.8-2.6x, respectively, a single node of Polaris. The applications were scaled up to 512 nodes showing good scaling efficiency across the board. Finally, we discuss useful concepts and experiences gained in running diverse applications on diverse HPC systems.

Read the paper · More papers on PaperTik