A Study of the prefetcher impact on high-performance computing applications
Valéria Soldera Girelli · Lume (Universidade Federal do Rio Grande do Sul) · 2021
Data prefetching algorithms are widely used in modern processors as a tool to mitigate the higher latency of memory accesses with respect to processor latency to execute instructions. However, understanding the contribution of prefetching to the application performance is a difficult task when we consider the high complexity found in the several architectures and prefetchers available. Developing accurate architecture simulators is also a challenge, especially when considering High-Performance Computing systems (HPC) with several processor cores. In this work, we contribute to shed light on the role of data prefetchers in the performance of parallel HPC applications, considering both the prefetcher algorithms offered in the real hardware and in the simulators. We performed a careful experimental investigation, executing the NAS parallel benchmark (NPB) on a real Skylake machine and in a simulated environment with the ZSim and Sniper simulators, using prefetcher algorithms offered by both Skylake and the simulators. Our experimental results show that: (i) prefetching from the L3 to L2 cache is responsible for the larger percentage of performance improvement, (ii) the memory contention in the parallel execution constrains the effectiveness of the prefetcher, (iii) the parallel memory contention in Skylake is poorly simulated by ZSim and Sniper, and (iv) the non-inclusive L3 cache present in the Skylake architecture hinders the accurate simulation of NPB with the Sniper prefetchers.