Architectural and operating system support for inexpensive, efficient shared memory
Leonidas I. Kontothanassis · 1996
Shared memory provides an attractive and intuitive programming model for large-scale parallel computing, but requires a coherence mechanism to allow caching for performance while ensuring that processors do not use stale data in their computation. Implementation options range from distributed shared memory emulations on networks of workstations to tightly-coupled, hardware-coherent, distributed shared memory multiprocessors. Previous work indicates that performance varies dramatically from one end of this spectrum to the other. Hardware cache coherence is fast, but also costly and time-consuming to design and implement, while DSM systems provide acceptable performance on only a limited class of applications. This dissertation claims that an intermediate hardware option--memory-mapped network interfaces that support a global physical address space, without cache coherence--can provide most of the performance benefits of fully cache-coherent hardware, at a fraction of the cost. Such network interfaces can be found both in tightly-coupled single-chassis multiprocessors and more loosely connected networks of workstations. We use the term NCC-NUMA (Non Cache Coherent, Non Uniform Memory Access) for machines with a global physical address space but no cache coherence. As part of the CASHMERe project we have developed a family of coherence protocols that take full advantage of the hardware properties of NCC-NUMA machines. These protocols implement a variant of lazy release consistency, allowing multiple concurrent writers per coherence block. However the use of the global physical address space makes it possible to improve protocol performance in several crucial ways. In particular it provides cheap access to directory data structures, efficient synchronization, and an inexpensive solution--writing through to a unique memory location--to the problem of merging inconsistent pages. Results indicate that these protocols running on NCC-NUMA hardware provide performance improvements over DSM systems by as much as an order of magnitude, and approach the performance of full hardware coherence. The dissertation also examines the impact of architectural trends, such as latency and bandwidth, on the relative performance of DSM, NCC-NUMA, and CC-NUMA hardware. NCC-NUMA systems remain very close to the knee of the price performance curve for a wide variety of architectural parameters. Finally, the dissertation examines the performance impact of adding hardware asynchrony and fine-grain access control to our family of protocols. For machines with flexible coherence maintenance mechanisms and fine-grain access control, our family of protocols provides consistently better performance than the current protocol of choice (eager release consistency).