Performance of CUDA-based FX cross correlation implementations
Xinping Deng, Yang Wu, Shengpu Niu, Yujing Zhang, Nan Zhao, Qing Shi, Haobao Shi · 2024
FX Cross correlation is a computation and memory intensive task; it is normally a bottle neck of signal processing in real time scenarios. As GPU is designed to process data in parallel, we decided to implement the algorithm on GPU for a better performance (comparing with CPU implementations). In this paper, we present 4 CUDA-based cross correlation implementations. The initial version did not perform very well. We then optimized it with share memory on GPU and improved its performance by a factor of 4. We then realized that we could get a better performance by doing cross correlation with optimized matrix multiplication CUDA libraries. In the end, we built two cross correlation pipelines with selected libraries (xGPU and tensor core) and compared their performance with our optimized one. We found out that these pipelines are much faster (the tensor core-based implementation is about 10 times faster) than our optimized implementation.