A performance evaluation of optimal hybrid cache coherency protocols
Jack E. Veenstra, Robert J. Fowler · 1992
The caches within a multiprocessor typically use either a write-invalidate protocol or a write-update protocol to maintain consistency. The recently introduced MIPS R4000 processor allows operating system software to select, on a per-page basis, which multiprocessor cache coherence protocol (write-invalidate versus write-update) the hardware will use. The availability of the R4000 and the prospect of even more flexible hardware motivated us to examine the potential performance advantages of allowing user-level control over the choice of coherence protocol on a per-page basis and to ask whether more powerful hybrid protocols provide substantially more benefit. We examine the potential benefits of three classes of hybrid protocols: (1) hybrid protocols that choose statically, at the beginning of the program, between write-invalidate (WI) or write-update (WU) on a per-page basis, (2) hybrid protocols that choose statically between WI or WU for each cache block, and (3) dynamic hybrid protocols that can choose between WI or WU at each write. In order to determine how much potential benefit could be obtained by each of these protocol classes, we used trace-driven simulations to evaluate the optimal off-line protocol for each class. We found that the use of a hybrid protocol can substantially reduce the cost of memory references for most of the programs studied. A few programs can also realize large additional benefits from a per-block static hybrid protocol compared to a per-page static hybrid protocol. None of the programs, however, receive a significant additional benefit from using a dynamic hybrid protocol compared to the per-block static hybrid protocol, unless cache block sizes are larger than 16 words (64 bytes).