Neat SIMD: Elegant vectorization in C++ by using specialized templates
Matthias Gross · 2016
Most of today's processors provide strong support for single instruction multiple data (SIMD, or vector) operations. But because modern compilers often fail to generate optimal SIMD code, explicitly written SIMD code is still the most performant. However, native SIMD code has some major disadvantages, mainly that it is quite cumbersome, that it is not easily ported to different instruction sets, and that typically multiple versions of similar code have to be written; e.g. a scalar, a streaming SIMD extensions (SSE) and an advanced vector extensions (AVX) version. Neat SIMD resolves all these disadvantages by using specialized templates in C++ that enable one neat version of easily portable code to be used for both scalar and SIMD instructions without sacrificing any performance in comparison to native SIMD code. Neat SIMD has no dependencies to any external library nor is it restricted to any specific instruction set or to a newer C++ version. The examples shown here use SSE and AVX. This publication is written such that all crucial code elements and implementation hints are given here. An interested software engineer can use those code elements, reconstruct the remaining code with ease, and start to use effective Neat SIMD code for own purposes.