The Use of Instruction-Based Prediction in Hardware Shared- Memory
Stefanos Kaxiras · Minds at UW (University of Wisconsin) · 1998
In this paper we propose Instruction-based Prediction as a means to optimize directory-based cache coherent NUMA shared-memory. Instruction-based prediction is based on observing the behavior of load and store instructions in relation to coherent events and predicting their future behavior. Although this technique is well estab- lished in the uniprocessor world it has not been widely applied for optimizing transparent shared-memory where pre- diction —in the form of adaptive cache coherence protocols— is typically address-based. The advantage of this technique is that it requires very few hardware resources in the form of very small prediction tables per node. In con- trast, address-based prediction typically requires storage proportional to the memory and/or cache size. To show the potential of instruction-based prediction we propose and evaluate four different optimizations: i) a migratory sharing optimization, ii) a wide sharing optimization iii) a pairwise sharing optimization, and iv) a producer-consumer opti- mization based on speculative execution. With execution-driven simulation and a set of ten benchmarks we show that: i) for the first two optimizations, instruction-based prediction performs comparably to and in some cases outperforms address-based schemes while never using more than 72 (5-byte) entries in any node's prediction table; ii) for pair- wise sharing there is no significant benefit over the default pairwise optimization of our base protocol. Finally we provide evidence that the producer-consumer optimization based on speculative execution can yield performance improvements.