SIMDization of Small Tensor Multiplication Kernels for Wide SIMD Vector Processors

Christopher I. Rodrigues, Amarin Phaosawasdi, Peng Wu · 2018

Developers often rely on automatic vectorization to speed up fine-grained data-parallel code. However, for loop nests where the loops are shorter than the processor's SIMD width, automatic vectorization performs poorly. Vectorizers attempt to vectorize a single short loop, using (at best) a fraction of the processor's SIMD capacity. It is not straightforward to vectorize multiple nested loops together because they typically have memory accesses with multiple strides, which conventional methods cannot profitably vectorize.

Read the paper · More papers on PaperTik