Adding low-cost hardware barrier support to small commodity clusters

Torsten Hoefler, Torsten Mehlan, Frank Mietke, Wolfgang Rehm · ARCS Workshops · 2006

The performance of the barrier operation can be crucial for many parallel codes. Especially distributed shared memory systems have to synchronize frequently to ensure the proper ordering of memory accesses. The barrier operation is often performed on top of point-to-point messages and the best algorithm scales with O(log2P · L) in the LogP model. We propose a cheap hardware extension which is able to perform the task of synchronization in nearly constant time and implement a driver inside the Open MPI framework to speedup the MPI Barrier() call. We test our implementation with the parallel implementation of Abinit and the MPI overhead decreases by nearly 32%.

Read the paper · More papers on PaperTik