Experiment management support for parallel performance tuning
Karen L. Karavanic, Barton P. Miller · 2000
The development of a high-performance parallel system or application is an evolutionary process. It may begin with models or simulations, followed by an initial implementation of the program. The code is then incrementally modified, and continues to evolve throughout the applications's lifespan. At each step, a key question for developers is: how and how much did the performance change? This question arises while comparing an implementation to models or simulations; considering versions of an implementation that use a different algorithm, communication or numeric library, or language; studying code behavior by varying number or type of processors, type of network, type of processes, input data set or work load, or scheduling algorithm; and benchmarking or regression testing. We present a design and prototype implementation of an experiment management environment designed to answer performance questions that span multiple program executions from all stages of the lifespan of an application. We have developed a concise representation for the set of executions collected over the life of an application. In our model, information from all experiments for one application, including the components of the code executed, execution environment, and performance data collected, is gathered in the Program Space. We developed techniques for automating comparison between measured executions. The structural difference operator determines differences in the source code and the runtime environment; the performance difference operator compares performance results and reports results that differ by more than a specified amount. We present several case studies exploring the use of these operators with large- scale parallel applications. We also developed a novel approach to automated performance diagnosis that uses application data gathered in previous executions to guide the search for performance bottlenecks. Adding historical knowledge about an application provides a means for the tool to perform more effective diagnosis. We evaluated our technique using different versions of an MPI application on an IBM SP/2, and found reductions of 31% to 98% in the time needed to locate performance bottlenecks.