A framework for hierarchical single-copy MPI collectives on multicore nodes

George Katevenis, Manolis Ploumidis, Manolis Marazakis · 2022

Collective operations are widely used by MPI applications to realize their communication patterns. Their efficiency is crucial for both performance and scalability of parallel applications. For deriving efficient MPI implementations, significant effort is put to keep pace with advances and capabilities of the underlying hardware and interconnect. Recent processor advances have led to nodes with higher core counts and complex internal structures and memory hierarchies. Such nodes are able to host tens to hundreds of processes and thus, performance of MPI collectives at the intra-node level becomes critical. In this work, we propose a framework for collective operations at the intra-node level, that aims to lower latency and increase bandwidth. Our approach utilizes knowledge of internal node structure to construct hierarchical algorithms, and XPMEM to achieve single-copy transfers. Pipelining is used to overlap communication at different levels of the hierarchy. We evaluate the proposed approach through several microbenchmarks and real-world MPI applications. For evaluation purposes, we compare the proposed approach with implementations of similar schemes from two recent studies. Our evaluation with microbenchmarks for Broadcast and Allreduce shows speedup up to 2.$5x$and$3x_{2}$respectively, over UCC and OpenMPI's default collectives implementation. Compared to recent research studies, we improve Broadcast by up to$5x_{2}$and Allreduce by up to$7x$. We reduce the time of three applications PiSvM, miniAMR and CNTK, by up to 12%, 52% and 12%, respectively, over the next best-performing alternative.

Read the paper · More papers on PaperTik