Memory Efficiency via Offloading in Warehouse-Scale Datacenters
Parthasarathy Ranganathan · Communications of the ACM · 2025
Memory Efficiency via Offloading in Warehouse-Scale DatacentersLarge warehouse-scale computers (WSCs) underpin all the cloud computing services we use daily-whether it is Web search, video streaming, social networks, or even emerging AI chatbots or agents.The memory subsystem in these computers poses one of the biggest challenges in their design and operation: Across the industry, Big Tech companies such as Amazon, Google, Meta, and Microsoft spend billions of dollars buying memory and consume hundreds of megawatts powering them.Sadly, this problem is only getting worse, exacerbated by slowing of technology scaling trends (like Moore's law) and exploding demand for more data and correspondingly more memory-for example, artificial intelligence (AI) workloads.One approach to address the costs of memory is to use tiers.Most workloads have a working dataset that includes both hot (more frequently used) and cold (less frequently used) data.Assigning the hot data to the fast but expensive memory while moving the colder data to a less expensive, albeit slower memory tier can make a dramatic difference to the total spending on memory.There are multiple ways to create this second, less expensive memory tier.We could use less costly, slower storage media, such as Flash and solid-state devices, or we could use faster memory and compress the data there.However, to make this work, we have to answer some important questions.How do we move data between the different memory tiers so that the applications do not slow down too much?This is particularly challenging since, in the same way memory dominates costs, memory also dominates performance in WSCs.Even relatively small variations in average memory latencies can lead to large performance slowdowns that can wash out any cost savings.There are other challenges too: How do we do move data transparently so that we do not need to change the plethora of workloads that run on these large cloud computing systems?And, most importantly, how do we do all this, at warehouse-scale, across heterogeneous workloads and diverse memory technologies?The accompanying paper, "TMO: Transparent Memory Offloading in Datacenters," addresses these questions.Focusing on an application-transparent, kernel-driven approach, TMO seeks to answer the two underlying questions of memory-tier management to meet the previously mentioned objectives: when to offload (and how much), and what memory to offload.The paper shows how its answers to these questions were very effective in the context of a real-world at-scale deployment at Meta.For the first question, the paper introduces a new metric called pressure stall information (PSI), tracked at the kernel level.Unlike prior approaches that use proxy metrics, such as page fault