DESA: Dataflow Efficient Systolic Array for Acceleration of Transformers

Zhican Wang, Hongxiang Fan, Guanghui He · IEEE Transactions on Computers · 2025

Transformers have become prevalent in various Artificial Intelligence (AI) applications, spanning natural language processing to computer vision. Owing to their suboptimal performance on general-purpose platforms, various domain-specific accelerators that explore and utilize the model sparsity have been developed. Instead, we conduct a quantitative analysis of Transformers. (Transformers can be categorized into three types: Encoder-Only, Decoder-Only, and Encoder-Decoder. This paper focuses on Encoder-Only Transformers.) to identify key inefficiencies and adopt dataflow optimization to address them. These inefficiencies arise from1)diverse matrix multiplication,2)multi-phase non-linear operations and their dependencies, and3)heavy memory requirements. We introduce a novel dataflow design to support decoupling with latency hiding, effectively reducing the dependencies and addressing the performance bottlenecks of nonlinear operations. To enable fully fused attention computation, we propose practical tiling and mapping strategies to sustain high throughput and notably decrease memory requirements from$O(N^{2}H)$to$O(N)$. A hybrid buffer-level reuse strategy is also introduced to enhance utilization and diminish off-chip access. Based on these optimizations, we propose a novel systolic array design, named DESA, with three innovations:1)A reconfigurable vector processing unit (VPU) and immediate processing units (IPUs) that can be seamlessly fused within the systolic array to support various normalization, post-processing, and transposition operations with efficient latency hiding.2)A hybrid stationary systolic array that improves the compute and memory efficiency for matrix multiplications with diverse operational intensity and characteristics.3)A novel tile fusion processing that efficiently addresses the low utilization issue in the conventional systolic array during the data setup and offloading. Across various benchmarks, extensive experiments demonstrate that DESA archives$5.0\boldsymbol{\times\thicksim}8.3\boldsymbol{\times}$energy saving over 3090 GPU and$25.6\boldsymbol{\times\thicksim}88.4\boldsymbol{\times}$than Intel 6226R CPU. Compared to the SOTA designs, DESA achieves$11.6\boldsymbol{\times\thicksim}15.0\boldsymbol{\times}$speedup and up to$2.3\times$energy saving over the SOTA accelerators.

Read the paper · More papers on PaperTik