A GPU Implementation of GPBiCGSafe-Algorithm-Based Navier-Stokes Solvers for 3D Incompressible Flows A GPU Implementation of 3D Navier-Stokes Solvers
Huynh Quang Huy Viet, Hiroshi Suito · 2014
an N × N dimensional matrix A can be seen as an one-dimensional array of length N which is stored in the global memory of a GPU. The vectors b and the result vector Ab of dimension N are also stored in the global memory. Zeroes are appended or prepended to the bands depending on the position of the band in the banded matrix A. Figure 2 (bottom) shows a process of multiplication of a tri-banded matrix A with a vector b. Each band of the matrix A is appended or prepended with zeroes corresponding the position of the band and then multiplied with the vector b which is shifted to left or right similarly. The products are summed to obtain the result vector Ab.