Intel ® in-Memory Analytics Accelerator: Performance Characterization and Guidelines

Jaeyoung Kang, Qirong Xia, Ipoom Jeong, Yongjoo Park, Nam Sung Kim · 2025

Improvements in CPU performance have significantly slowed due to the demise of Dennard scaling, making it increasingly challenging to cost-effectively process the exponentially growing volumes of data. As a result, there is a growing trend of offloading frequently used functions to hardware accelerators to enhance application performance while reducing expensive CPU cycle consumption. In line with this trend, Intel introduced the InMemory Analytics Accelerator (IAA) as an on-chip accelerator in its$\mathbf{4}^{\text{th}}$-generation Xeon®Scalable CPUs (Sapphire Rapids). IAA is designed to offload common big data and in-memory analytics functions from CPUs, such as CRC64, expand, extract, scan, and select, in addition to (de)compression. As IAA is integrated with CPUs as an on-chip accelerator, it can directly access the CPU's cache and memory in a cache-coherent manner, offering low latency and power consumption with reduced programming complexity compared to off-chip accelerators. In this paper, we first introduce the hardware and software architectures of IAA, highlighting its latest features. We then evaluate IAA's performance using various microbenchmarks and widely used analytics applications, including Pandas, Citus, and ClickHouse. Finally, based on our findings, we provide guidelines for effectively utilizing and optimizing IAA for analytics applications.

Read the paper · More papers on PaperTik