Collective Reduction Operation on Cray X1 and Other Platforms

Rolf Rabenseifner, Panagiotis A. Adamidis · 2004

A 5-year-proling in production mode at the University of Stuttgart has shown that more than 40% of the execution time of Mes- sage Passing Interface (MPI) routines is spent in the collective commu- nication routines MPI Allreduce and MPI Reduce. Although MPI im- plementations are now available for about 10 years and all vendors are committed to this Message Passing Interface standard, the vendors' and publicly available reduction algorithms could be accelerated with new al- gorithms by a factor between 3 (IBM, sum) and 100 (Cray T3E, maxloc) for long vectors. This paper presents ve algorithms optimized for dif- ferent choices of vector size and number of processes. The focus is on bandwidth dominated protocols for power-of-two and non-power-of-two number of processes, optimizing the load balance in communication and computation. The new algorithms are compared also on the Cray X1 with the current development version of Cray's MPI library (mpt.2.4.0.0.13)

Read the paper · More papers on PaperTik