Large-scale FFTs and convolutions on Apple hardware
Richard E. Crandall, J. Klivington, D. A. Mitchell · 2008
Impressive FFT performance for large signal lengths can be achieved via a matrix paradigm that exploits the modern concepts of cache, memory, and multicore/multithreading. Each of the large-scale FFT implementations we report herein is built hierarchically on very fast FFTs from the standard OS X Accelerate library. (The hierarchical ideas should apply equally well for low-level FFTs of, say, the OpenCL/GPU variety.) By building on such established, packaged, small-length FFTs, one can achieve on a single Apple machine—and even for signal lengths into the billions—sustained processing rates in the multi-gigaflop/s region.