Characterizing Load Imbalance in Real-World Networked Caches
Qi ping Huang, Helga Gudmundsdottir, Ýmir Vigfússon, Daniel A. Freedman, Ken Birman, Robbert van Renesse · 2014
Modern Web services rely extensively upon a tier of in-memory caches to reduce request latencies and alleviate load on backend servers. Within a given cache, items are typically partitioned across cache servers via consistent hashing, with the goal of balancing the number of items maintained by each cache server. Effects of consistent hashing vary by associated hashing function and partitioning ratio. Most real-world workloads are also skewed, with some items significantly more popular than others. Inefficiency in addressing both issues can create an imbalance in cache-server loads.