iWAPT 2018 Invited Speaker 2

David E. Tanner · 2018

Messy, complicated and demanding. Enormous difficult-to-realize potential. Programming GPUs is like parenting a 3-year old. The Tensile project seeks to fully automate achieving peakperformance for all GEMMs, tensor contractions and convolutions; for all precisions, transposes, sizes and strides; and for all GPUs. Doing so requires addressing a variety of obstacles such as problem types, GPU performance bottlenecks, skinny matrices, small matrices, slow transposes, edge cases and emitting source and assembly kernels, auto-tuning kernels per problem size and auto-generating host library code. Tensile addresses these issues to achieve high-performing GEMMs and tensor contractions across a wide range of traditionally difficult problem sizes.

Read the paper · More papers on PaperTik