A Processing-using-Memory Architecture for Commodity DRAM Devices with Enhanced Compatibility and Reliability

Hoon Shin, Rihae Park, Jae Wook Lee · 2024

Processing-in-memory (PIM) paradigm is a promising solution to minimize the cost of data movement in the von Neumann architecture. Many proposals have explored possibilities to realize this paradigm in the main memory leveraging DRAM technologies. DRAM-based PIM technology can be implemented in two ways: processing-in-memory with digital compute units in peripheral area and processing-using-memory (PUM). Although computing within memory cells without digital compute units at peripherals is certainly appealing, DRAM-based PUM architectures still have various limitations to overcome. In particular, we identify three major challenges to bring PUM closer to production: compatibility with commodity DRAM microarchitecture, reliability, and coverage of operations. To address these challenges, we propose PRADA, a Processing-using-Memory Architecture based on DRAM cell Array. Unlike existing proposals, PRADA does not introduce any changes to the highly optimized cell area by not utilizing either designated rows for computation or dual-contact cells (DCCs) to implement NOT operation. Instead, PRADA proposes two new states in the bitline sense amplifier (BLSA) to implement NOT operation without any new additional circuitry. We also demonstrate how to implement various operations in PRADA, not only bit-wise logical operations but also both integer (INT) and floating-point (FP) arithmetic operations, while ensuring reliable bitline (BL) sensing. Compared to state-of-the-art PUM architectures, PRADA demonstrates 2.67--4.79× higher throughput for 8-bit integer multiply. For vector-ADD, PRADA achieves 3.09--3.13× speedups over the baseline, outperforming the other PUM architectures which show 1.04--2.07× speedups, while maintaining superior compatibility and reliability.

Read the paper · More papers on PaperTik