Specialized Accelerator for Arbitrary-Dimension Matrix Operations in Deep Learning

Fei Liu, Zhouquan Liu, Jing Zhang, Mingche Lai, Hui Ling Guo, Libo Huang · 2025

This paper addresses the significant challenge of executing inference tasks involving General Matrix Multiplication (GEMM) in deep neural networks(DNN) on resource-constrained edge systems. While previous research has introduced various accelerator architectures for GEMM, these solutions primarily support matrix multiplications of varying dimensions at the software level rather than the hardware level. We propose an optimization of data reuse and memory access patterns through the integration of the outer product method. Building on this, we developed an optimized workload allocation strategy to further minimize data transfer overhead and balance computational resource utilization. Additionally, a heterogeneous hardware module named the Matrix Acceleration Engine (MAE) has been designed and integrated into the RISC-V processor to accelerate prevalent GEMM operations in deep neural networks. Compared to traditional instruction sets, the RV-MAE processor achieves a speedup of up to 194.74× in matrix operations and up to 20.29× in convolutional neural network tasks. Compared to advanced matrix accelerators, the performance is improved by 5.33×. Furthermore, the MAE, evaluated in a 28nm CMOS process, exhibits an area of 9742.18µm2and a power consumption of 2.9mW, thus meeting the low-power demands of edge computing.

Read the paper · More papers on PaperTik