Instrumentation for distributed computing systems
David W. Krumme, George V. Cybenko, Alva L. Couch · Fall joint computer conference · 1987
At present there seems to be no shortage of workable strategies for measuring the performance of parallel computing hardware. Software simulators and software and hardware probes are available which can generate reams of printout about any algorithm or hardware configuration. But invariably the user cannot afford to read the reams of data produced. Thus the usual feedback the user gets is a few averages. While something can be learned from these averages, they in no way portray the totality of what is happening inside the machine. The problem with measurement tools is not that they cannot measure enough, but that the measurements are not made meaningful to the human user.Thus the real problem facing instrumentation specialists is not one of engineering, but rather of mathematics: each dataset has the potential to be so large and detailed that new theory is needed to accurately portray the complexities of computations without resorting to averages alone. We need a new methodology that allows the specialized hypotheses of parallel computation to be analyzed. Central to this new theory is the issue of causality. When testing an algorithm, one wants to know not just how well the algorithm performs but also which part of the algorithm caused this performance, i.e. what is the weakest part of the algorithm. Under parallel computing conditions such a weakness cannot always be isolated with confidence. Then the user is faced with testing hypotheses about what part of the algorithm causes poor performance. Hypothesis testing of this sort, concerning the relations between source code and performance, should be automated. Likewise, the hardware designer must often try to determine what part of the hardware is causing a particular high-level performance bottleneck. While this may be a harder problem than the relationship between source and performance, some parts of it should be automated as well.But even if we successfully develop an automated hypothesis testing instrument, we will not have solved the basic problem underlying parallel program development: there are too many hypotheses. How does one choose which hypothesis to test? Clearly this is not the domain of the instrument; there are exponentially many possible hypotheses and any algorithm for selecting one would sooner or later have to resort to exhaustive search. Further, there is an intimate connection between hypothesis selection and pattern recognition in the measured information. Thus the human mind, not the instrument, is most well suited for this activity. But current instruments produce information unsuitable for this type of analysis; there is too much data for the human user to scan.This is the motivation for the SEECUBE system display performance data for the hypercube multiprocessor. SEECUBE provides a variety of ways of looking at the data, and it supports the simultaneous display of multiple views so that different aspects of the same computation can be followed together. Current work on SEECUBE involves devising new kinds of windows into the data and developing the associated software probes, so that more types of hypotheses can be tested.It seems that the pattern recognition and hypothesis testing phases could be integrated into a single instrument. With such an instrument, specification of the hypotheses could be simplified by the pattern viewing structures, and testing could proceed at a phenomenal rate. Such an instrument will be required for the next generation of parallel machines: as systems are scaled up in size and complexity, the interrelationships will become even more deeply embedded and the user will not be able to economically perform algorithm analysis without assistance.