Parallel solution and optimization of microphysics module in GRAPES physics process

Aokang Xu, Jinfang Jia · 2025

Global and Regional Assimilation and Prediction System (GRAPES) is a domestically developed numerical weather prediction system in China. The microphysical process is the key physical process of cloud formation and precipitation and is also one of the hot spots in the calculation of the whole model. Despite the parallel processing capabilities of MPI and OpenMP, the computation time of the microphysics process remains relatively long. In order to achieve finergrained parallelization, the code is ported to Graphics Processing Unit (GPU) using a combination of memory coalescing and array scalarization methods. In addition, to address the issue of underutilization of GPU resources when multiple processes share a single GPU, the Compute Unified Device Architecture (CUDA) Multi-Processing Service (MPS) is introduced in the microphysics module to improve GPU utilization. Experimental results show that for a resolution of 50km, excluding data transfer time, the speedup achieved by using 4 GPUs in conjunction with 36 CPU cores reaches 11.59, and the speedup excluding data transfer time is 18.59. It has been verified that the parallel computation of the microphysics process on GPUs is stable and correct.

Read the paper · More papers on PaperTik