Hypercube Communication Performance
Carolyn L. McCreary, M. E. McArdle, J. D. McCreary · 1991
Performance information is essential to the design of efficient parallel programs. Whether the programmer has total control over determining the parallel execution of processes or whether some automatic means are used to parallelize serial or parallel code, the cost of the overhead due to parallelization must be known. This is especially important in systems where the costs of such mechanisms are high. In distributed systems, message passing is the primary means of communication between parallel processes, and message passing tends to be quite expensive. Users of such systems wrestle with a variety of questions related to performance issues. How is the data to be distributed among the processors? What grain size (amount of code executed serially on a processor) is required for efficient processing? How should the processes be assigned to the system nodes? Approximate answers to these questions can be given only if the performance characteristics of the system are known. This paper presents a step towards the development of a performance model for two distributed memory multiprocessor systems, the Intel iPSC/2 and iPSC/860 hypercubes. From actual timing measurements, we developed formulas to estimate the communication costs on the four possibilities for message passing among multiple processors: a single sender to a single receiver, single sender to multiple receivers, multiple senders to a single receiver, and multiple senders to multiple receivers. A practical method for synchronizing timing measurements among independently clocked processing nodes was developed to do this analysis.