A DRAM-based Near-Memory Architecture for Accelerated and Energy-Efficient Execution of Transformers
Gian Singh, Sarma B. K. Vrudhula · 2024
Transformers-based language models have achieved remarkable accuracy in various NLP tasks, employing self-attention mechanisms primarily based on matrix multiplication. However, their significant size leads to data movement issues, causing latency and energy efficiency challenges in conventional Von-Neumann systems. To mitigate these issues, several in-memory and near-memory architectures have been proposed. This paper introduces PACT-3D, a near-memory architecture featuring novel computing units integrated with DRAM banks. PACT-3D significantly reduces latency by 1.7 × and improves energy efficiency by 18.7 × compared to state-of-the-art near-memory architectures.