A Design of 16TOPS Efficient GEMM Module in Deep Learning Accelerator
Guoning Lu, Dong Yu Xu, Ning Wang, Xiao Li Zhang, Degen Zhen, Hong Lei, Yunlong Bai, Dehui Kong, Hang Ruan, Zhifeng Chi, Xiankui Xiong, Ke Xu · 2020
An efficient GEMM (general matrix multiplication) module is presented as a key computation unit in DLA (Deep Learning Accelerator). The core is located in a system which consists of a RISC-V processor for scalar operation and instruction distribution, and PE (Processing Engine) for vector operation and tensor operation. This paper focuses on the tensor operation module GEMM, which provides a highly-parallel processing capability through innovative data preprocessing and memory layout. To provide enough flexibility for various neural networks, the GEMM also supports multiple instructions dispatched from RISC-V processor. Meanwhile, hierarchical instruction queue ensures the scalability. The proposed design achieves 16 TOPS INT8 at 1GHz when implemented with TSMC 12nm technology.