Demanding Parallel FFTs: Slabs & Rods
Ian Kirker · 2008
Fourier transforms of multidimensional data are an important component of many scientific codes, and thus the efficient parallelisation of these transforms is key in obtaining high performance on large numbers of processors. For the three-dimensional case, two common decompositions of input data can be employed to perform this parallelisation – one-dimensional (slab), and two-dimensional (rod) processor grid divisions. In this report, we demonstrate implementation of the three-dimensional FFT routine in parallel, using component routines from a library capable of performing one-dimensional FFTs, and briefly examine considerations of creating a portable application which can use multiple libraries in C. We then examine the performance of the two decompositions on seven different highperformance platforms – HECToR, HPCx, Ness, BlueSky, MareNostrum, HLRB II, and Eddie, using each of the FFT libraries available on each of these platforms. Efficient scaling is demonstrated up to more than a thousand processors for sufficiently large data cube sizes, and results obtained indicate that the slab decomposition generally obtains greater performance. Additionally, our results indicate that FFT libraries installed on platforms we tested are largely comparable in performance, whether vendor-provided or open-source, except for the esoteric Blue Gene/L architecture, upon which ESSL proved to be superior.