Rootless Operations for MPI (RLO) v0.1
Tonglin Li, Quincey Koziol · OSTI OAI (U.S. Department of Energy Office of Scientific and Technical Information) · 2019
Flexible communication routines that enable high-performance computing (HPC) applications to operate with less synchronization are a great benefit to application algorithm creation and developer productivity. Further, communication operations that reduce synchronization in HPC applications while continuing to scale well are critical to application performance in the exascale era. We have implemented mechanisms that allow an application to scalably execute bcast/iAllReduce from any MPI rank without specifying the origin of the message in advance, demonstrating application development flexibility, highly scalable operation, and reduced application synchronization.