FlexGEMM: A Flexible Micro-kernel Generation Framework

Shunhong Wang · 2024

Deep learning (DL) has progressed by leaps and bounds in the past decade. Due to the immense computational cost of DL workloads, industry and academia have developed DL libraries with highly-specialized kernels for certain workloads/architectures, leading to numerous complex code bases that strive for performance. However, they are hard to maintain and not universal. In this paper, we propose a general-purpose, performance-portable micro-kernel FlexGEMM based on DPC++ SYCL implementation. By using parametric adjustment, FlexGEMM can easily adapt to different workloads and heterogeneous devices like GPUs and CPUs, and keep high-performance in the meantime. Task mapping schemes improve the cache and atomic efficiency. And if the workload and GPU meet the conditions, the data-racefree writing method can improve the GEMM performance by taking advantage of parallelism as much as possible. On Intel® Iris® Xe MAX Graphics (96EU), comparing FlexGEMM with the state-of-the-art library oneMKL, we demonstrate that our performance is on par with, and in some cases even surpasses it. In the case of tall-and-skinny matrices, FlexGEMM can achieve 2× to 20× speedup compared to oneMKL.

Read the paper · More papers on PaperTik