Rewriting and Optimizing Vector Length Agnostic Intrinsics from Arm SVE to RVV
Jhih-Kuan Lin, Yu-Lun Yang, Hung-Ming Lai, Jenq‐Kuen Lee · 2024
Advanced processors incorporate SIMD extensions to execute data-parallel operations efficiently. As technology advances, new generations of SIMD extensions evolve with longer vector register lengths, making target-specific program non-portable. To address this issue, architectures with vector length agnostic (VLA) programming models have emerged, which can scale across implementations with varying register lengths. However, utilizing SIMD hardware commonly involves programming with target-specific intrinsics. Despite VLA’s scalability within the same ISA, intrinsics programs designed for one VLA target would still encounter portability issues when deployed on other VLA architectures. Although automatic rewriting techniques exist, rewriting strategies for VLA architectures are a relatively new area of research with limited studies available. In this work, we present our rewriting strategies based on the open-source intrinsics rewriting library, SIMD Everywhere (SIMDe), for porting Arm SVE intrinsics to RISC-V Vector Extension (RVV). Our method efficiently transforms masks for predicated instructions between SVE and RVV formats. Additionally, we introduce algorithms for removing redundant mask computation. In our experiment, we evaluate our approach using compute kernels collected from the Arm SVE programming example document that are written in SVE intrinsics. Compared to the scalar implementations provided in the document, which are further vectorized by the compiler, our rewritten RVV still achieves speedup ranging from 1.39× to 28.9× in terms of dynamic instruction count.