AttBind: Memory-Efficient Acceleration for Long-Range Attention Using Vector-Derived Symbolic Binding

Weihong Xu, Jaeyoung Kang, Tajana Rosing · 2024

Transformer models have achieved a number of breakthrough results in a variety of complex tasks. Transformer's promising performance originates from multi-head attention (MHA), which can model long-range sequence data dependency. Better performance has been demonstrated to be obtained by increasing the sequence length$N$. However, scaling up the sequence length is extremely challenging for memory-constrained hardware because the naive Transformer requires quadratic$O(N^{2})$complexity. In this work, we address this challenge by leveraging the binding operation in vector symbolic architecture (VSA). We propose the memory-efficient MHA algorithm to simplify the MHA computation at the cost of linear complexity. Then, we present the ASIC hardware architecture with optimized timing and dataflow to accelerate the proposed algorithm. We extensively evaluate our design across various long-range attention tasks. Our experiments show that the accuracy is competitive to state-of-the-art MHA optimization approaches with lower memory consumption and inference latency. The proposed algorithm achieves 7.8× speedup and 4.5× reduction in data movement over the naive Transformer on ASIC. Meanwhile, our design supports 8 to 16 × sequence lengths compared to existing hardware accelerators.

Read the paper · More papers on PaperTik