From SMPs to FPGAs: Multi-Target Data-Parallel Programming
Barry Bond, Andrew A. Davidson, Lubomir Litchev, Satnam Singh · 2010
Is it possible to devise a data-parallel programming model that is sufficiently abstract to permit automatic compilation from a single description into very different execution targets like multicore processors, GPUs and FPGAs? And if such a compilation process is possible what is the cost of such an approach compared to the alternative of designing specific specialized implementations for each target? This paper examines a high level data-parallel programming model which is designed to target multiple architectures. Using this model we describe the implementation of the core component of a software defined radio system and its automatic compilation into SSE3 vector instructions running on a SMP system; an implementation for GPUs that targets both NVidia and ATI graphics cards; and an implementation that targets Xilinx FPGAs. For the GPU case we compare the performance and design effort of our Accelerator based system with several hand coded CUDA implementations. Our goal here is to qualitatively evaluate the amount of programming effort versus achieved performance in order to assess the viability of improving programmer productivity by raising the level of abstraction of the data-parallel programming model.