Cache coherence directories for scalable multiprocessors
Jr. Richard Thomas Simoni · 1992
this memory bandwidth problem, since multiple processors may be referencing the same memory modules. Furthermore, it is impossible to physically locate all memory nearby all of the processors, 2 CHAPTER 1. so some references must incur long access latencies due to interconnect delays. Average memory latency can be reduced somewhat by distributing the global shared memory across the processing nodes. Even so, if a processor is actively sharing data, then it will reference some data that is not locally resident. One approach to improving the characteristics of the memory system is to follow the lead of uniprocessor designers by pairing each processor with a cache. Caches improve latency and bandwidth by using small, fast memories that are tightly coupled with the CPU. In a multiprocessor, they reduce the bandwidth strain due to data sharing by allowing each processor to cache its own copy of a shared data value. Unfortunately, allowing multiple processors to simultaneously cache a given datum leads to the cache consistency problem, also known as the cache coherence problem. If one processor writes a shared data value in its cache, the other cached copies of the data become stale. A processor that reads a stale copy of the data does not receive the most recently written value, violating the shared memory model we would like to provide. The simplest solution to the cache coherence problem is to disallow the caching of shared data. The performance effects stemming from the longer latencies that result from caching only private data are demonstrated by Figure 1.1. The vertical axis shows the fraction of uniprocessor utilization that is achieved by each processor. The horizontal axis shows the fraction of data references that are to shared data. The solid curve indicates the...