Locality-Aware Process Mapping for High Performance Collective MPI-IO on FEFS with Tofu Interconnect
Yuichi Tsujita, Atsushi Hori, Yutaka Ishikawa · 2014
Collective MPI-IO for non-contiguous access has been frequently used in not only its direct MPI-IO application programming interface (API) calls, but also scientific application driven parallel I/O libraries such as the HDF5, which utilizes MPI-IO APIs underneath its parallel I/O APIs. Since performance improvements to collective MPI-IO is a key issue in parallel I/O operations, we have been focusing significant efforts in that area. In this paper, we propose an alternative locality-aware process mapping optimization scheme in two-phase I/O in ROMIO, which is a representative MPI-IO library to achieve higher collective I/O performance during non-contiguous accesses. Sequential process mapping, which choose processes in order when accessing a target file system in a parallel manner, does not always result in I/O performance improvements. In our study, we implemented a locality-aware process mapping approach suitable for the Fujitsu Exabyte File System (FEFS) with the Tofu Interconnect in the K computer. Through performance evaluations, we were able to achieve I/O throughput improvements of up to 50% relative to the original library performance levels.