Optimization of Triton matrix multiplication for tensor core

Yixuan Gao, Jianan Li, Xiaonan Chai, Junxiang Wang, Dong Zhang · 2025

Matrix multiplication has important and widespread applications in the field of computer science, serving as the foundation of high-performance computing systems and the core of deep learning model applications. Tensor Core is a hardware unit designed specifically for accelerating matrix multiplication, which can greatly improve the computational efficiency of matrix multiplication. Triton is a flexible and efficient deep learning compiler that simplifies the operator development process while providing efficient training and inference performance. However, its block configuration can affect the computational efficiency of the operator. When a certain dimension of the block configuration is small, it is difficult for memory access and computation to completely overlap in multi-level pipeline optimization, resulting in Tensor Core being unable to fully utilize parallel computing performance. In response to this issue, dual chain splitting optimization was designed and developed to adjust the memory access and computation instructions while improving the shared memory access efficiency. The experiment tested different data storage modes and operand shapes. The average acceleration ratio under different data storage modes is 1.3, and the utilization rate of Tensor Core increased by an average of 4.74%. The average acceleration ratio under different operand shapes is 1.21, and the utilization rate of Tensor Core increased by an average of 3.75%.

Read the paper · More papers on PaperTik