Membrane: Accelerating Database Analytics with DRAM-Based PIM Filtering and Schema Denormalization

Akhil Shekar, Kevin P. Gaffney, Martin Prammer, Khyati Kiyawat, Lingxi Wu, Helena Caminal, Ziyao Fan, Yimin Gao, Ashish Venkat, José F. Martínez, Jignesh M. Patel, Kevin Skadron · ACM Transactions on Architecture and Code Optimization · 2025

In-memory database query processing frequently involves substantial data transfers between the CPU and memory, leading to inefficiencies due to the Von Neumann bottleneck. Processing-in-Memory (PIM) architectures offer a viable solution to alleviate this bottleneck. In our study, we employ a commonly used software approach that streamlines JOIN operations into simpler selection or filtering tasks via pre-join denormalization, thereby making the query processing workload more amenable to PIM acceleration. This research explores the DRAM design landscape to evaluate how effectively these filtering tasks can be executed efficiently across the DRAM hierarchy and their effect on overall application speedup. We also find that operations such as aggregates are better executed on the CPU than on PIM. Thus, we propose a cooperative query processing framework that capitalizes on both CPU and PIM strengths, where (i) the DRAM-based PIM block, with its massive parallelism, supports scan operations while (ii) CPU, with its flexible architecture, supports the rest of the query execution. This allows us to utilize both PIM and CPU where appropriate and prevent dramatic changes to the overall system architecture. With these minor modifications to the system architecture and a customized version of the DuckDB database to integrate offloaded scan operations into the CPU-side processing, our methodology enables accurate end-to-end performance evaluations using established analytical benchmarks such as TPC-H and the Star Schema Benchmark (SSB). Our findings show that this novel mapping approach improves performance, delivering a \(5.92x/6.5x\) speedup (TPCH/SSB) compared to a traditional schema and \(3.03-4.05x\) speedup compared to a denormalized schema with \(9-17\%\) memory overhead, depending on the degree of partial denormalization. Further, we provide insights into query selectivity, memory overheads, and software optimizations in the context of PIM-based filtering, which better explain the behavior and performance of these systems across the benchmarks.

Read the paper · More papers on PaperTik