HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented Generation

Chaoqiang Liu, Haifeng Liu, Dan Chen, Yu Huang, Yi Zhang, Wenjing Xiao, Xiaofei Liao, Hai Jin · 2025

By integrating external knowledge bases, Retrieval-augmented Generation (RAG) enhances natural language generation for knowledgeintensive scenarios and specialized domains, producing content that is both more informative and personalized.RAG systems typically consist of two fundamental stages: retrieval and generation.The retrieval stage experiences low bandwidth utilization due to its random and irregular memory access patterns.Meanwhile, the generation stage is also constrained by memory bandwidth limitations, which arise from involving a significant number of General Matrix-Vector Multiplications (GEMV) operations.These two stages collectively lead to memory bottlenecks within RAG systems.Recent efforts leverage HBM-based Processing-in-Memory (PIM) to accelerate conventional Large Language Models (LLMs).However, the retrieval stage incurs substantial storage overhead due to the need to maintain large-scale knowledge bases, resulting in a capacity bottleneck.Solely relying on HBM-based PIM in RAG is both costly and insufficient to meet the capacity demands.Fortunately, DIMM-based PIM provides a low-cost, high-capacity alternative that complements HBM.In this work, we propose HeterRAG, a novel heterogeneous PIM acceleration system for RAG.It combines

Read the paper · More papers on PaperTik