A comparison of parallelization techniques for irregular reductions

Hwansoo Han, Chau‐Wen Tseng · 2001

A large class of scientific applications are comprised of irregular reductions on large data sets. On shared-memory multiprocessors these reductions are typically parallelized by computing partial results into replicated buffers, then combining the values into shared data using synchronization. Recently, a number of alternative techniques have been developed based on selective privatization, local writes, and synchronized writes. In this paper, we present a more efficient version of the local write algorithm which is 56% faster on average. We then experimentally compare the performance of each technique using a number of representative kernels. Results show speedups vary greatly depending on application characteristics such as connectivity, locality, and adaptivity. In general, we find the local write technique provides the best performance, particularly when applications display good locality.

Read the paper · More papers on PaperTik