swYAKL: A Data Parallel Runtime on Manycore Architecture

Yanghui Ye, Junshi Chen, Hong Qian, Kunxian Lin, Yuanhang Li, Hong An · 2024

The traditional architecture of supercomputers comprises a control unit and heterogeneous accelerators with distinct memory spaces, necessitating developers to code separately for each accelerator. The notion of performance portability has been introduced to facilitate the swift adaptation of a singular codebase across multiple heterogeneous accelerators, enabling the same code to efficiently harness the performance of various platforms with minimal alterations. Frameworks such as Kokkos, RAJA, and YAKL accomplish this through a data parallel model, primarily targeting SIMT devices but also applicable to SIMD devices. With the advent of manycore architectures that provide high levels of parallelism and performance, it becomes imperative to extend performance portability to these architectures, which currently require users to manually partition and map tasks to fully exploit their capabilities. This paper introduces a data parallel programming runtime for the Sunway manycore architecture and adapts YAKL to this platform, termed swYAKL. This runtime capitalizes on the computational characteristics of Sunway, including task division and mapping, and supports Sunway’s CPE LDM and native vectorization. It employs a specialized task division method for multi-dimensional stencil kernels to leverage the DMA capabilities of Sunway’s CPE. Performance evaluations using several common computational kernels and a proxy application named miniWeather indicate that swYAKL achieves speedups exceeding 100x compared to execution on Sunway’s MPE for various workloads, and under specific conditions, it surpasses the performance of the NVIDIA A100 GPU.

Read the paper · More papers on PaperTik