GPU-Ready GASNet Implementation on the TCA Proprietary Interconnect Architecture

Kenta Sato, Norihisa Fujita, Toshihiro Hanawa, Taisuke Boku, Khaled Z. Ibrahim · 2016

TCA (Tightly Coupled Accelerators) is a concept of direct interconnection among GPUs over multiple nodes, which has an experimental implementation named PEACH2. Since PEACH2 is a special device architecture for realizing low-latency GPU-to-GPU communication, the programming cost to port GPU-ready parallel code such as CUDA+MPI is a crucial problem. In this paper, we extend GASNet, a widely used communication layer for PGAS languages, to become GPU ready. This layer is an extension of the GASNet runtime system for CPU-GPU and GPU-GPU communication as well as for CPU-CPU communication. The block-stride transfer function of PEACH2 enabled us to achieve communication performance that is up to 2.1x higher than that obtained with the InfiniBand implementation without this feature.

Read the paper · More papers on PaperTik