Architectural and compiler support to hide coherence misses in distributed shared-memory multiprocessors
Josep Torrellas, David A. Koufaty · 1997
As the gap between processor and memory speed increases, using powerful processors in shared-memory multiprocessors will not be productive if they often stall waiting for the memory system to supply data and instructions. Moreover, as technology advances, data coherence-induced misses will become more important due to the larger and more integrated local memories that will make conflict misses relatively inexpensive. We make several contributions to cope with the problem of coherence-induced misses. First, we propose four producer-initiated data forwarding primitives designed to reduce the impact of coherence-induced misses. The primitives try to hide the latency by enabling the producers to send data to consumers right at the time that the data is written. Second, we present a compiler algorithm that replaces some of the regular writes by one of the proposed primitives and evaluate its performance effect. Finally, we compare our primitives with data prefetching and propose new techniques to integrate them. Using execution-driven simulations, we found that our algorithm for data forwarding reduces the execution time by an average of 30%-40%, depending on the size of the local memory hierarchy. Several optimizations helped increase the performance. In particular, we found that forwards should always be delayed and combined, both locally and in the directory. Also, forwards do not need to either update the memory of the home node or send the data all the way up to the second level cache. Our evaluation of data prefetching resulted in performance improvements of about 30%, regardless of the memory size. For the applications studied, while for the the average improvement of data forwarding is smaller than in data forwarding, neither technique outperforms the other in all cases. The proposed integrated techniques are able to further improve on the performance of either case, resulting in speedups of 48% on average, regardless of the size of the local memories.