Scalable Performance Analysis Methods for the Next Generation of Supercomputers

Felix Wolf, Daniel F Becker, Markus Geimer, Brian J. N. Wylie · JuSER (Forschungszentrum Jülich) · 2008

Facing increasing power dissipation and little instruction-level parallelism left to exploit, computer architects are realizing further performance gains by using larger numbers of moderately fast processor cores rather than by further increasing the speed of uni-processors.As a consequence, supercomputing applications are required to harness much higher degrees of parallelism in order to satisfy their growing demand for computing power.However, writing code that runs efficiently on large numbers of processors remains a significant challenge.To address this challenge, the Helmholtz-University Young Investigators Group Performance Analysis of Parallel Programs at the Jülich Supercomputing Centre develops performanceanalysis tools to diagnose inefficiencies in supercomputer applications and works with application developers to analyze and improve the performance of their codes.In this contribution, we highlight the research activities of our group during the past two years and give an outlook on future work.At the centre of our report lies the development of SCALASCA, a performanceanalysis tool that has been specifically designed for large-scale systems and that allows the automatic identification of harmful wait states in applications running on thousands of processors.

Read the paper · More papers on PaperTik