GRONA : A Framework for Gather-and-Reduce On Near-Memory Accelerators
Aman Kumar Sinha, Pei-Yi Liu, Yuhao Fang, Jhih-Yong Mai, Bo-Cheng Charles Lai · 2023
Gather-and-reduce (GnR) is a collective operation widely used in Big-data analytic workloads. Memory-bound fine-grained data-accesses from massive lookup-tables during Gather, followed by compute-bound arithmetic and logical processing during Reduce, are suitable for low-latency computations on emerging Near-DRAM Processing (NDP) architectures. While NDP on general Dual Inline Memory Modules (DIMMs) offer simple integration and find commercial adoption, cross-level data-sharing over standard buses remains a bottleneck. Maintaining standard functionality of the DIMM require careful schemes for scheduling, coordination and memory management in order to achieve the potential throughput.We propose GRONA (Gather-and-Reduce On Near-Memory Accelerators), a framework for intra-DIMM throughput-oriented reprogrammable Gather-and-Reduce. GRONA adopts a multilevel reprogrammable PE platform equipped with small caches to support all GnR applications. GRONA framework achieves effective decoupling of bank-group and rank-level data accesses and GnR execution, and displayed up-to 2.9x performance gains for applications such as FM-Index query and SparseLengthSums compared to NDP state-of-the-arts.