The performance and scalability of distributed shared memory cache coherence protocols

John LeRoy Hennessy, Mark A. Heinrich · 1998

Distributed shared memory (DSM) machines are becoming an increasingly popular way to increase parallelism beyond the limits of bus-based symmetric multiprocessors (SMPs). The cache coherence protocol is an integral component of the memory system of these DSM machines, yet the choice of cache coherence protocol is often made based on implementation ease rather than performance. The Stanford FLASH (FLexible Architecture for SHared memory) multiprocessor provides an environment for running different cache coherence protocols on the same underlying hardware. In the FLASH system, the cache coherence protocols are written in software that runs on a programmable controller specialized to efficiently run protocol code sequences. Within this environment it is possible to hold everything else constant and change only the cache coherence protocol code that the controller is running, thereby making visible the impact that the protocol has on overall system performance. This dissertation examines the performance of four full-fledged cache coherence protocols for the Stanford FLASH multiprocessor. The four protocols are a simple bit-vector/coarse-vector protocol, the dynamic pointer allocation protocol, the IEEE standard Scalable Coherent Interface (SCI) protocol, and a flat Cache Only Memory Architecture (COMA-F) protocol. The protocols are compared when running a mix of scalable scientific applications at different machine sizes from 1 to 128 processors. In addition, results are also shown for less-tuned versions of each application, as well as for systems with smaller processor caches. The results show that cache coherence protocol performance can be critical in DSM systems, with over 2.5 times performance difference between the best and worst protocol in some configurations. In addition, no single existing protocol always achieves the best overall performance. Surprisingly, the best performing protocol changes with machine size—even within the same application! In the end, the results argue for programmable protocols on scalable machines, or a new and more flexible cache coherence protocol. For designers who want a single architecture to span machine sizes and cache configurations with robust performance across a wide spectrum of applications using existing cache coherence protocols, flexibility in the choice of cache coherence protocol is vital.

Read the paper · More papers on PaperTik