Cache memory design and performance issues in shared-memory multiprocessors
Farnaz Mounes-Toussi · 1996
The widening gap between the speed of processors and main memory has encourage the use of cache memories to reduce the average memory access time by exploiting the spatial and temporal locality of memory references in a program. In shared-memory multiprocessors, however, data sharing reduces the effectiveness of cache memories by introducing the cache coherence problem and the inherent coherence overhead. The many solutions to the cache coherence problem attempt to reduce the coherence overhead by using different mechanism to enforce coherence, to detect accesses to stale data, and to reduce the impact of false-sharing. In this thesis, we examine the performance effect of these different mechanisms. We discuss how the coherence overhead depends on the sharing characteristics of a program and define a model for characterizing the sharing behavior of parallel programs. We also propose and evaluate a partially-valid write-through write-merged coherence mechanism to tolerate false-sharing. This mechanism by itself does not entirely eliminate redundant write-through, however, so we also define a compiler algorithms to eliminate some of the redundant write-through to memory. Using simulations and our model, we identify the high processor-memory network traffic as the main problem associated with an updating coherence enforcement strategy, and the high number of misses as the main problem associated with an invalidating coherence enforcement strategy. We also show that the severe performance penalty resulting from inaccurate compile-time analysis in a software-only coherence detection mechanism favors the use of a hardware-only or a combined hardware-software mechanism. Our study of false-sharing indicates that the performance penalty of false-sharing is severe when a program has a large amount of sharing and a fine granularity of sharing. Our studies of a partially-valid write-through cache with and without a merging write-buffer emphasize the difficulty of using a hardware mechanism to adjust to the programs' sharing characteristics. Consequently, we address this difficulty using compile-time analysis in conjunction with a hardware mechanism to enhance performance.