Toward efficient resource utilization of a GPU-accelerated AI supercomputer
Rohin Arora · 2021
Many production high-performance computing (HPC) systems have integrated multiple accelerators on the same node to better serve machine learning and data-intensive workloads. To improve the design and operations of such heterogeneous accelerator-based supercomputers, we performed a detailed analysis of system operations, job characteristics, user behavior, and trend forecasting on MIT Supercloud (ranked in the top 50 in the Top500 supercomputer list). Our analysis reveals that HPC users are increasingly using supercomputing clusters for interactive and development jobs. We identified a novel kind of jobs, referred to as "exploratory jobs", which consume a significant portion of computing resources -- an emerging trend that is not commonplace on many supercomputers yet, but expected to grow in the future. Subsequently, we developed several strategies to improve the resource utilization and user experience by targeting such exploratory jobs.--Author's abstract