AiMX: Accelerator-in Memory Based Accelerator for Cost-effective Large Language Model Inference (Invited)

Haerang Choi, Guhyun Kim, Woojae Shin, Jongsoon Won, Changhyun Kim, Hyunha Joo, Byeongju An, Gyeongcheol Shin, Jeongbin Kim, Dayeon Yun, Jae‐Han Park, Yosub Song, Byeongsu Yang, H.-H.S. Lee, Seungyeong Park, Wonjun Lee, Seonghun Kim, Yonghoon Park, Yousub Jung, Il Kon Kim · 2024

We presented an Accelerator-in-Memory (AiM) device and AiM-based LLM inference acceleration system. LLM inference can be divided into prompt phase and response phase. Considering the characteristics of LLM inference, we proposed a disaggregated inference system where the prompt phase is executed on high-throughput GPUs or NPUs, and the response phase is executed on AiM. Using AiM for single GEMV operations can ideally achieve up to 16 times the performance. The measured performance of the prototype AiM-based Accelerator is 1.7 times higher than that of a comparable GPU, and the expected performance at the highest data rate is 11.7 times higher.

Read the paper · More papers on PaperTik