Automatic Generation of Micro-kernels for Performance Portability of Matrix Multiplication on RISC-V Vector Processors
Francisco D. Igual, Luís Piñuel, Sandra Catalán, Hèctor Martínez, Adrián Castelló, Enrique S. Quintana–Ort́ı · 2023
In this paper, we propose and evaluate several optimized implementations of the general matrix multiplication (gemm) on two different RISC-V architecture cores implementing the RISC-V vector extension (RVV): C906 and C910 from T-HEAD. Specifically, we address the performance portability problem across these processor cores by means of an automatic assembly code generator, written in Python, capable of emitting RVV code for high performance computing (HPC), with a variety of combinations of specific and general optimizations.