Latency-Aware and Auto-Migrating Page Tables for ARM NUMA Servers
Hongliang Qu, Peng Wang · Electronics · 2025
The non-uniform memory access (NUMA) architecture is the de facto norm in modern server processors. Applications running on NUMA processors may suffer significant performance degradation (NUMA effect) due to the non-uniform memory accesses, including data and page table accesses. Recent studies show that the NUMA effect of long-running memory-intensive workloads can be mitigated by replicating or migrating page tables to nodes that require accesses to remote page tables. However, this technique cannot adapt to the situation where other applications compete for the memory controller. Furthermore, it was only implemented on x86 processors and cannot be readily applied on ARM server processors, which are becoming increasingly popular. To address this issue, we designed the page table access latency aware (PTL-aware) page table auto-migration (Auto-PTM) mechanism. Then we implemented it on Linux ARM64 (the Linux kernel name for AArch64) by identifying the differences between the ARM architecture and the x86 architecture in terms of page table structure and the implementation of the Linux kernel source code. We evaluate it on real ARM NUMA servers. The experimental results demonstrate that, compared to the state-of-the-art PTM mechanism, our PTL-aware mechanism significantly enhances the performance of workloads in various scenarios (e.g., GUPS by 3.53x, XSBench by 1.77x, Hashjoin by 1.68x).