Energy Efficiency and Performance of Auto-Vectorized Loops on Intel Xeon Processors
Olga V. Moldovanova, Mikhail G. Kurnosov, Aleksey Mel'nikov · 2018 3rd Russian-Pacific Conference on Computer Technology and Applications (RPC) · 2018
This paper evaluates how well modern compilers Intel C/C++, GCC C/C++, LLVM/Clang and PGI C/C++ auto-vectorize loops. We use the Extended Test Suite for Vectorizing Compilers (ETSVC) as a benchmark. We estimate time, energy, power and speedup by running the loops in scalar and vector modes for different data types (double, float, int, short int) and determine loop classes which the compilers used in the paper fail to vectorize. Our study shows that the compilers evaluated could vectorize 39-77% of the total number of loops in the ETSVC package. The best results were shown by the Intel C/C++ Compiler, and the worst ones - by the LLVM/Clang compiler. The compilers failed to vectorize loops containing conditional and unconditional branches, function calls, induction variables, variable loop bounds and iteration count, as well as such idioms as 1st or 2nd order recurrences, search loops and loop rerolling. Our experimental platform was the dual CPU system (NUMA server, 2 x Intel Xeon E5-2620v4, Intel Broadwell microarchitecture) with the Intel Xeon Phi 3120A co-processor. We use the Running Average Limit Power (RAPL) Machine Specific Registers (MSRs) to obtain the energy measurements.