Accelerator Design using 3D Stacked Capacitorless DRAM for Large Language Models

Janak Sharda, Po-Kai Hsu, Shimeng Yu · 2024

Large language models (LLMs) have been immensely useful for natural language processing tasks. However, the current model sizes are increasing exponentially, along with generating large amounts of intermediate data. Here, we propose to use the capacitorless 3D stackable DRAM, which is an emerging memory enabling scaling of DRAM in the vertical direction like 3D NAND Flash. A 3D DRAM can store much larger LLMs compared to conventional DRAM at higher density. Further, to reduce the intermediate data size, we propose to use a layer-wise sparsity-quantization hybrid (LSQH) algorithm, which induces sparsity based on calculations performed using low-bit quantization to reduce both the energy consumption and the data storage requirements. Finally, a 3D heterogeneously integrated accelerator is designed by stacking a 3D DRAM with logic dies designed in the 3 nm technology node, which exploits the LSQH algorithm for the Llama2 model. The evaluation of the proposed system shows an energy efficiency of > 25 TOPS/W and an area efficiency of > 14.4 TOPS/mm2, with minimal drop in accuracy. Further model scaling is performed to obtain the energy efficiency and area efficiency for larger LLMs.

Read the paper · More papers on PaperTik