Accelerating Confidential Recommendation Model Inference With Near-Memory Processing

Wenjie Xiong, Liu Ke, Maxim Ostapenko, Yong-Min Tai, Yeongon Cho, Joonho Song, Jinin So, Kyung-Soo Kim, Yongsuk Kwon, Jin Jung, Jieun Lee, Byeongho Kim, Shin-haeng Kang, Sukhan Lee, Jeong‐Hyeon Cho, Kyomin Sohn, Xuan Zhang, Hsien-Hsin S. Lee, G. Edward Suh · IEEE Transactions on Dependable and Secure Computing · 2025

Trusted Executing Environments (TEEs) in hardware designs protect program execution from other untrusted software programs in the processor as well as untrusted off-chip hardware components. Meanwhile, Near-Memory Processing (NMP) has shown performance and energy benefits on memory-intensive workloads. Recently, novel memory encryption schemes have been proposed to allow TEEs to leverage the benefits of NMP without requiring trust in the NMP components. In this paper, we present a system design of confidential computing with NMP that can be directly used in Intel SGX, a TEE platform available in commercial processors today. We develop the full software stack and evaluate the results on commercial processors with the emulated AxDIMM, an FPGA-based NMP platform. In our case study on personalized Deep Learning Recommendation Model (DLRM) inference, the proposed confidential computing in NMP achieves up to 1.51× latency reduction and up to 2.57× throughput improvement.

Read the paper · More papers on PaperTik