Multiple endpoints for improved MPI performance on a lattice QCD code
Larry Meadows, Ken-Ichi Ishikawa, Taisuke Boku, Masashi Horikoshi · 2018
This paper provides results using multiple threads and a high-performance MPI implementation of MPI_THREAD_MULTIPLE applied to a Lattice QCD Code (CCS-QCD) and run on the Oakforest-PACS machine. Performance has improved from the baseline code by as much as 1.8x for smaller lattice sizes.