AttenTPU: Tensor Processor for Attention Mechanism with Fine-Grained Padding
Zhihao Du, Yike Li, Chao Chen, Zheng Wang · 2023
Transformer-based models have achieved state-of-the art performance in many Artificial Intelligence (AI) tasks. The core component of the transformer is the attention mechanism, which includes computation-intensive operators-such as MatMul and Linear-raising demand for hardware support. Many Institute, and researchers have proposed their hardware acceleration methods in recent years. However, most of these works mainly focus on algorithm optimization and hardware performance. Little attention has been paid to the variable sequence lengths in the attention mechanism. In the hardware design, the variability of sequence length brings a challenge to the load-store unit and on-chip data movement. In this work, we first propose a hardware-friendly fine-grained padding method aiming to handle the variability of sequence length in scaled dot-product attention. Then, a hardware architecture named AttenTPU is proposed using the padding method. The experimental results indicate that compared to the CPU and GPU platform, our accelerator achieves speed-ups of$3.43\times$and$1.41\times$, respectively.