Collaborative development of a performance analysis methodology

Flynn, Thomas, Weinzierl, Tobias · Zenodo (CERN European Organization for Nuclear Research) · 2025

Use cases of computationally intensive resources, such as High-Performance Computing (HPC), are increasingly diversifying with the advent of machine learning and the rise of the digital arts and humanities, while traditional HPC continues to push the boundaries and ask for more capability. This means we have a more demanding and diverse audience of HPC developers, and we therefore want to build a community on robust methods and practices for HPC use. Performance analysis must play an integral role in how developers use their software, because performant software allows researchers to do more research, but also to responsibly use the shared HPC resources economically and ecologically. This talk reports on findings from a series of workshops and investigative works undertaken at Durham University to develop a performance analysis methodology that is broad but nonetheless rigorous. Typically, performance analysis training begins with profiling and tracing tools. However, this can often leave developers using tools to generate performance data but without a strategy of how to systematically analyse this data or navigate through the space of analysis options. We propose a ‘methodology first’ approach, in which our use of tools is informed by our methodology, i.e., we want to plan our analysis and reach for the right tool for the job. It also comprises a clear roadmap of how to start an analysis and what steps to follow one by one. This methodology at the highest level is based on five performance topics: core, intra-node, inter-node, I/O and GPU. The work to develop this methodology has engaged HPC-focused software engineers to collaboratively design decision trees and metrics to probe these individual performance topics. One outlook of this work is to engage with broader communities to understand how these methodologies can be adapted to incorporate the latest HPC workflows including machine learning. A recording of this session is available on YouTube: https://youtu.be/pRwqPY8fBjo

Read the paper · More papers on PaperTik