Exploring GEMM Operations on Different Configurations of the Gemmini Accelerator
Dennis Agyemanh Nana Gookyi, Eunchong Lee, Kyungho Kim, Sung‐Joon Jang, Sang-Seol Lee · 2022 19th International SoC Design Conference (ISOCC) · 2022
The General Matrix to Matrix Multiplication (GEMM) has recently been employed to improve Deep Neural Networks (DNN) performance. This paper explores GEMM operations of different matrix dimensions and dataflow options including weight stationery (WS) and output stationery (OS) on various systolic array sizes (8×8, 16×16, 32×32, and 64×64) of the Gemmini full-stack open-source DNN accelerator generator. The GEMM operation performance in a baseline Rocket CPU core is compared with two Gemmini configurations: with and without a hardware image to column (im2col) block. The different configurations with different systolic array sizes are synthesized using an FPGA device and compared in terms of hardware area, maximum frequency, and vector-less power consumption. Overall, a Gemmini configuration with a 16×16 array, a hardware im2col module, and a WS dataflow recorded the highest performance per area of 1.04×, 2.31×, and 8.84× compared with 8×8, 32×32, and 64×64 array configurations respectively.