Remote Memory Prefetching: Is Coarse-grained Fine?
James McMahon, Vinita Pawar, Ryan Stutsman · 2025
CXL raises new questions about tiering, pooling, and remote memory access. In most disaggregated memory approaches, compute nodes access remote memory pools at page (4~KB) granularity. This is the case for two reasons: to help amortize high remote access costs and because the host CPUs' address translation hardware is set up for 4~KB pages. While fetching and caching whole remote pages helps with applications that have high spatial locality, for some applications it can introduce contention for cache capacity since potentially cold or unrelated adjacent data is also cached. Additionally, this can increase network bandwidth utilization.