An Implementation of Tensor Product Patch Smoothers on GPUs
Cu Cui, Paul Grosse-Bley, Guido Kanschat, Robert Strzodka · SIAM Journal on Scientific Computing · 2025
Abstract. We present a GPU implementation of vertex-patch smoothers for higher-order finite element methods in two and three dimensions. Analysis shows that they are memory bound not with respect to GPU DRAM but with respect to on-chip scratchpad memory. Multigrid operations are optimized through localization and reorganized local operations in on-chip memory, achieving minimal global data transfer and a conflict-free memory access pattern. Performance tests demonstrate that the optimized kernel is at least two times faster than the straightforward implementation for the Poisson problem across various polynomial degrees in two and three dimensions, achieving up to 36% of the peak performance in both single and double precision on an NVIDIA A100 GPU.