Modeling communication overhead: MPI and MPL performance on the IBM SP2

Zhiwei Xu, Kai Hwang · IEEE Parallel & Distributed Technology Systems & Applications · 1996

The authors use timing experiments on the IBM SP2 to develop an overhead-quantifying method for evaluating communication performance on message-passing multicomputers. Massively parallel processor (MPP) designers and users can apply this method to reveal architectural bottlenecks and to trade off between computations and communications for parallel applications optimization. This article presents a systematic method to estimate the communication overheads of three message-passing operations (point-to-point communication, collective communication and collective computation) on MPPs. We validated this method by a performance study of the Message-Passing Interface (MPI) and the IBM Message-Passing Library (MPL) on the IBM SP2 at the Maui High-Performance Computing Center. We measured overheads of the three communication operations for various combinations of machine size and message length. We used the collected timing data to derive the overhead expressions. For a given MPP, the timing measurements need to be performed only once.

Read the paper · More papers on PaperTik