Comparative benchmarking: matrix multiplication on a multicore coprocessor and a GPU
Maryam Salim, Ali O. Akkirman, Mert Hidayetoğlu, Levent Gürel · 2015
This paper reports the performances of an Intel Xeon Phi coprocessor and an Nvidia Tesla GPU for multiplication of large matrices. For this purpose, various libraries, such as Intel MKL and MAGMA, are employed with different execution modes of the coprocessor. We compare the performances of the coprocessor and the GPU in terms of running time, memory requirement, and programming difficulty for the special case of matrix-matrix multiplication.