Locality-improved FFT implementation on a graphics processor

Sergio Romero, Marı́a A. Trenas, Eladio Gutiérrez, Emilio L. Zapata · 2007

Abstract: The growing computational power of modern graphics processing units is making them very suitable for general purpose computing. These commodity processors operate generally as parallel SIMD platforms and, among other factors, the effectiveness of the codes is subject to a right exploitation of the underlying memory hierarchy. This paper deals with the implementation of the Fast Fourier Transform on a novel graphics architecture offered recently by NVIDIA. Such an implementation takes into consideration memory reference locality issues, that are crucial when pursuing a high degree of parallelism, that is, a good occupancy of the processing elements. The proposed implementation has been tested and compared to the manufacturer’s own implementation.

Read the paper · More papers on PaperTik