Performance Evaluation of Loop Body Splitting for Fast Modal Filtering in SCALE-DG on A64FX
Xuanzhengbo Ren, Yuta Kawai, Hirofumi Tomita, Seiya Nishizawa, Takahiro Katagiri, Tetsuya Hoshino, Daichi Mukunoki, Masatoshi Kawai, Toru Nagai · 2025
Modern general-purpose central processing units (CPUs) benefit from the integration of Single Instruction, Multiple Data (SIMD) architectures.Scalable Vector Extensions (SVE) is one of Arm's SIMD architectures designed for HPC.Fujitsu's A64FX is the first Armbased processor to incorporate hardware-implemented SVE alongside High-Bandwidth Memory 2 (HBM2).However, in the A64FX, the high latencies of SIMD instructions, such as fused multiply-add (FMA), combined with the limited capacity of reservation stations, can result in inefficient Out-of-Order (OoO) execution.SCALE-DG is an atmospheric dynamical core that uses the discontinuous Galerkin Method (DGM).Modal filtering, an essential procedure in SCALE-DG, has an optimized version called fast modal filtering, which suffers from the OoO execution issue due to dot products with long vector lengths in the loop body.This issue can be alleviated by splitting the loop body into multiple parts.However, since the vector length of the dot product varies with the polynomial order (𝑃) in SCALE-DG, it remains unclear how the number of splits affects the performance of fast modal filtering.In this paper, we present an evaluation across 𝑃 values in the range of 3 to 11 with different splitting numbers and combinations.The results indicate that when 𝑃 ≤ 7, performance degraded in most cases, with only a few cases achieving positive speedups (1.01x to 1.02x) after splitting the loop body.For 8 ≤ 𝑃 ≤ 11, splitting had a consistently