Scalable Collation and Presentation of Call-Path Profile Data with CUBE
Markus Geimer, Björn Kuhlmann, Farzona Pulatova, Felix Wolf, Brian J. N. Wylie · JuSER (Forschungszentrum Jülich) · 2007
Developing performance-analysis tools for parallel applications running on thousands of processors is extremely challenging due to the vast amount of performance data generated, which may conflict with available processing capacity, memory limitations, and file system performance especially when large numbers of files have to be written simultaneously.In this article, we describe how the scalability of CUBE, a presentation component for call-path profiles in the SCALASCA toolkit, has been improved to more efficiently handle data sets from thousands of processes.First, the speed of writing suitable input data sets has been increased by eliminating the need to create large numbers of temporary files.Second, CUBE's capacity to hold and display data sets has been raised by shrinking their memory footprint.Third, after introducing a flexible client-server architecture, it is no longer necessary to move large data sets between the parallel machine where they have been created and the desktop system where they are displayed.Finally, CUBE's interactive response times have been reduced by optimizing the algorithms used to calculate aggregate metrics.All improvements are explained in detail and validated using experimental results.