Optimizing Collective Operations in Hybrid Applications
Aurèle Mahéo, Patrick Carribault, Marc Pérache, William Jalby · 2014
The advent of multicore and manycore processors in clusters advocates for combining MPI with a shared memory model like OpenMP in high-performance parallel applications. But exploiting hardware resources with such models can be sub optimal. Thus, one approach is to use the hybrid context to perform MPI communications. In this paper, we address this issue with a concept of hybrid collective communications, which consists in using OpenMP threads to parallelize MPI collectives. We validate our approach on several MPI libraries (IntelMPI and MPC), improving the overall time up to a factor of 5.29×, in a real world application.