Using Remote Accelerators to Improve the Performance of the FFTW Library
Santiago Mislata, Federico Silla · 2016
Hardware accelerators have forced a change in high performance computing. Their use has enabled an increment in the performance of data centers. For this reason, developers have decided to port many applications belonging to diverse science fields, such as biology or chemistry, to hardware accelerators like GPUs (Graphics Processing Units). Nevertheless, not all the applications have been able to be ported to accelerators, either due to the difficulty it entails or because the cost of performing such port does not outweigh the benefits. In this work we present a new middleware that uses remote accelerators to perform the intensive computations of scientific libraries, initially intended to be executed in the local CPUs. Forwarding the computationally intensive parts of these libraries to remote accelerators is done in a transparent way to applications, not having to modify their source code. This implementation of the new middleware focuses on offloading the FFTW library. An in-depth performance evaluation of the new middleware is presented, considering several node configurations and InfiniBand networks. NVIDIA Tesla K40 GPUs are used as accelerators. Results shown that the FFTW library experiences a speedup larger than 25x, despite having to transfer data back and forth to the remote server owning the accelerator.