A Framework for Near Memory Processing With Computation Offloading and Load Balancing

Satanu Maity, Manojit Ghose, Sudeep Pasricha · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2025

Due to the increasing demand for off-chip data transfers, the traditional Von-Neumann architecture faces challenges with modern data-intensive applications, leading to the memory-wall problem. Near-memory processing (NMP) provides a solution by placing computation units near the main memory, which reduces off-chip data transfers and improves system performance. Under this paradigm, some portions of the application are transferred and executed on the NMP side, known as computation offloading. This article introduces a novel computation offloading approach for NMP-enabled 3-D memory systems, considering several critical factors collectively, including data locality information at the last-level cache and execution time estimation of the offloadable portions, which have not been collectively explored in existing studies. Further, this article proposes two different load-balancing strategies to distribute workloads among the NMP cores in the 3-D memory, thereby improving overall performance further. Extensive experiments using a variety of applications from different application domains demonstrate the effectiveness of the proposed approach. Our approach achieves a maximum speedup of$2.34\times $and$1.91\times $compared to traditional and state-of-the-art approaches, respectively. The proposed approach also reduces off-chip data transfers by nearly$5.3\times $compared to traditional computing architectures. Furthermore, the best-proposed approach reduces energy consumption by 26% (maximum) and 21% (average) compared to various state-of-the-art approaches.

Read the paper · More papers on PaperTik