Parallel 5 point SOR for solving the Convection Diffusion equation using graphics processing units
Yiannis Cotronis, Elias Konstantinidis, Maria A. Louka, Nikolaos M. Missirlis · 2013
In this paper we study a parallel form of the SOR method for the numerical solution of the Convection Diffusion equation suitable for GPUs using CUDA. To exploit the parallelism offered by GPUs we consider the fine grain parallelism model. This is achieved by considering the local relaxation version of SOR. More specifically, we use SOR with red black ordering with two sets of parameters ωij and ω ′ ij for the 5 point stencil. The parameter ωij is associated with each red (i+j even) grid point (ij), whereas the parameter ω ′ ij is associated with each black (i+j odd) grid point (ij). The use of a parameter for each grid point avoids the global communication required in the adaptive determination of the best value of ω and also increases the convergence rate of the SOR method [2]. We present our strategy and the results of our effort to exploit the computational capabilities of GPUs under the CUDA environment. Additionally, a parallel program utilizing manual SSE2 (Streaming SIMD Extensions 2) vectorization for the CPU was developed as a performance reference. The optimizations applied on the GPU version were also considered for the CPU version. Significant performance improvement was achieved with the three developed GPU kernel variations.