MPI's Reduction Operations in Clustered Wide Area Systems.
Thilo Kielmann, Rutger F. H. Hofman, Henri E. Bal, Aske Plaat, Raoul A. F. Bhoedjang · Data Archiving and Networked Services (DANS) · 1999
The emergence of meta computers and computational grids makes it feasible to run parallel programs on large-scale, geographically distributed computer systems. Writing parallel applications for such systems is a challenging task which may require changes to the communication structure of the applications. MPI’s collective operations (such as broadcast and reduce) allow for some of these changes to be hidden from the applications programmer. We have developed MAGPIE, a library of collective communication operations optimized for wide area systems. MAGPIE’s algorithms are designed to send the minimal amount of data over the slow wide area links, and to only incur a single wide area latency. This paper discusses MPI’s collective reduction operations. Compared to systems that do not take the topology into account, such as MPICH, large performance improvements are possible. For larger messages, best performance is achieved when the reduction function is associative. On moderate cluster sizes, using a wide area latency of 10 millisecond and a bandwidth of 1 MByte/s, operations execute up to 8 times faster than MPICH; application kernels improve by up to a factor of 3. Due to the structure of our algorithms, the advantage increases for higher wide area latencies.