A Hardware Architecture for Shared Residuals and Simplified Symmetric-PWL-Based GELU
Zhelai Ding, Dezheng Zhang, Dong Wang · 2024
In this paper, we propose an innovative hardware acceleration scheme to optimize key components of Visual Transformer (ViT) and similar models. It features two main innovations. The first is a shared residual mechanism that integrates residual structures, class tokens, and position embeddings using a "first adding then concatenating" strategy to reduce duplicated hardware resource consumption. The second is a simplified piecewise linear (PWL) method for GELU that effectively lowers computational complexity. Validated on the AMD Alveo U55C platform, this approach reduces resource usage and latency. When combined with the shared residual kernel in our hardware accelerator, this approach achieves nearly double the performance compared to CPU computation on FPGA, paving the way for enhanced Transformer hardware acceleration.