Hardware Acceleration for Detrs Using Systolic Array Architecture

Xinchen Li, Min Xie, Ruixiao Zhao, Yujie Zhang · 2025

Transformer models have shown great potential on visual tasks in recent studies, especially the DETR series models, which have revolutionized the traditional convolutional neural network-based object detection tasks. However, the high computation complexity in DETRs poses great challenge in hardware deployment. To enable deployment in edge computing scenarios, efficient hardware acceleration is essential. Among the computations of DETRs, the Multi-Head Attention (MHA) mechanism is the most demanding in terms of hardware resources and computation time. In this paper, we present a hardware accelerator based on a systolic array architecture to compute the MHA module efficiently. A dedicated softmax accelerator is integrated within the array sequence to handle the softmax operation between matrix multiplications. We take RT-DETR, the latest model of DETRs as an example to accelerate its core AIFI MHA module. Our design is described by hardware description language and evaluated on a Xilinx xczu15eg FPGA. Experimental results show an$8.98 \times$speedup over CPUs and a$2.2 \times$speedup over GPUs, along with a$2.1 \times$improvement in energy efficiency compared to CPUs. With our accelerator, the difficulty of hardware deployment for DETRs in hardware-constrained environments should be greatly reduced.

Read the paper · More papers on PaperTik