A tensor core architecture for large-scale matrix multiplication operations on GPGPU processor
Zhengfu Feng, Ang Li, Dan Wu, Huan Zhang, Lidong Xing, Yu Liu · 2025
Recently, as DNN has demonstrated extraordinary performance and efficiency in areas such as image processing and target recognition, research on DNN (deep neural network) and its dedicated hardware accelerators has also increased, and the demand for computing power and versatility of neural networks has further expanded. Today’s DNN accelerators mainly focus on accelerating specific operators, such as GEMM (general matrix multiplication) and convolution operations. As the most mainstream computing chip at present, GPGPU provides powerful computing power support for many current computing tasks. However, the traditional computing units in GPGPU are not sufficient to meet the strong demand for computing power for deep neural networks. This paper proposes a tensor core based on the GPGPU processor to ensure that the GPGPU processor can provide strong support for large-scale matrix multiplication operations.