Design and application of cache coherent multiprocessors
Ashwini K. Nanda · 1993
With electronic device speed approaching physical limits, parallel processing has become the next logical step toward achieving higher computing speeds. Parallel processors fall mainly into two categories, namely, message passing and shared memory. Shared Memory multiprocessors offer a simple programming model and have gained widespread popularity for general purpose programming needs. When the number of processors in a multiprocessor system increases, the interconnection network becomes more complex, and so the memory latency time also increases. Private cache memories offer an effective solution to this problem. However, the use of private cache memories brings with it the associated cache coherence problem. Many cache coherence schemes have been proposed and used for the shared memory multiprocessors over the past decade. There are several design and performance issues involved in the use of cache coherence schemes in multiprocessor systems. In this dissertation, we address the following issues related to design and application of cache coherent shared memory multiprocessors: (1) Performance evaluation of cache coherent multiprocessors, (2) Design of efficient interconnection networks to support cache coherence protocols, (3) Specification and verification of cache coherence protocols and (4) Efficient mapping of applications on cache coherent multiprocessors. Analytical models are developed for various cache coherent multiprocessor systems and are validated using an event driven simulation model. Since the analytical models run at least two orders of magnitude faster than the simulation programs, and are reasonably accurate, we prefer them over the later as an evaluation tool. A new network, called the Multistage Bus Network(MBN) is designed to support cache coherence protocols efficiently. The MBN consists of multiple stages of buses connected in a way similar to the Multistage Interconnection Network(MIN). The MBN has the efficient broadcasting advantages of bus networks and scalable bandwidth advantages of the MIN's, and therefore is a suitable candidate for building large cache coherent multiprocessor systems. Directories of state information are introduced to the bus based MBN switches and a multi stage snoopy protocol is proposed to maintain cache coherence. Another scheme, called the MIN with Directories(MIND) scheme where directories are introduced into the crossbar based MIN switches, is also discussed. We show that the MBN scheme consistently outperforms the MIND scheme as well as the existing directory schemes. A formal technique based on Communicating Finite State Machines is proposed to specify and verify the cache coherence protocols. The technique is demonstrated using two bus based snoopy protocols. Finally, a technique is presented to estimate the communication cost of executing parallel programs in the presence of cache coherence protocols. The estimate of communication cost is used to efficiently map parallel applications on cache coherent multiprocessor systems. A Sequent Balance 8000 multiprocessor is used to verify the application mapping technique.