Performance Analysis of the NVIDIA HPC SDK and AMD AOCC Compilers in an HPC Cluster Using Pooled, Robust and Relative Metrics

Yectli A. Huerta · 2024

System characterization identifies which micro-architectural components are not being optimally utilized. CPUs are equipped with Performance Monitoring Units (PMUs) which make it possible to understand how efficiently different architectural components are used. The number of monitoring units and the granularity of data these units provide is CPU dependent. Performance variation makes it difficult to just report on a single metric that can be used to summarize complex systems and their corresponding results. The art of reporting performance results involves the use of plots, ratios of performance metrics, and summary statistics. Simple plots could suffice for stand alone systems that require no share resources such as filesystems. Some of these approaches, such as the use of simple plots, have their shortcomings, because in complex systems like HPC clusters, there are a number of anomalies that can have an effect in the variability of performance. Metrics, plots and summary statistics that don't account for variability and skewness of outliers will fail to provide a consistent and accurate picture of a system's performance. In this study, we propose an approach that uses pooled metrics, z-score normalization, metrics of dispersion and the purchasing parity index to quantify the relative differences between runs made on different computational nodes using two different compilers. Our goal is to make it possible to better interpret performance results through the use of robust techniques. Using our proposed process, we analyzed the NVIDIA HPC SDK and the AMD AOCC compilers. The AOCC compilers were optimized to support the AMD Zen architecture, while the NVIDIA compilers are widely used in platforms with NVIDIA GPU accelerators. We found that the NVIDIA compilers had more dispersion in the runtime results and that dispersion was generated by outliers, and that in ten out of the twelve tested benchmarks, the NVIDIA compiled binaries required less CPU cycles to achieve a comparable IPC rate when compared to the AOCC results.

Read the paper · More papers on PaperTik