Realizing Out-of-Core Stencil Computations Using Multi-tier Memory Hierarchy on GPGPU Clusters
Toshio Endo · 2016
The memory wall problem is one of major obstacles against the realization of extremely fast and large scale simulations. Stencil computations, which are important kernels for CFD simulations, have been highly successful on GPU clusters in speed, due to high memory bandwidth and computation speed of accelerators. However, their problem scales have been limited by small capacity of GPU device memory. In order to support larger domain sizes than not only device memory capacity but host memory, we extend our approach that combines locality improved stencil computations and a runtime library that harnesses memory hierarchy. This paper describes the extended version of HHRT library that supports multi-tier memory hierarchy, which consists of device memory, host memory and high speed flash SSD devices. And we demonstrate our approach effectively realizes out-of-core execution of stencil computations, whose problem scales are three time larger than host memory capacity.