Tensile: Auto-Tuning GEMM GPU Assembly for All Problem Sizes

David E. Tanner · 2018

Generic matrix matrix multiplication (GEMM) on graphics processors (GPU) has long been the target of both tuning to find a fastest kernel for a GPU and hand-written assembly to achieve highest possible efficiency. Tensile is an open source (https://github.com/ROCmSoftwarePlatform/Tensile) tool which builds on both of these, by auto generating assembly kernels, auto benchmarking the kernels against a range of problem sizes, then auto generating library C code which selects the optimal kernel for a given problem size. Not only can the kernels achieve near-peak efficiency for the best problem sizes, but the size-tuned kernels can achieve speedups of 2X to 350X compared to any single-tuned kernel optimized for a given GPU architecture.

Read the paper · More papers on PaperTik