Clustered distributed virtual shared memory for large-scale multiprocessing
Andrew Erlichson · 1998
One potentially attractive way to build large scale shared-memory machines is to use small-to-medium scale shared-memory machines as clusters, interconnecting the clusters with an off-the-shelf network. To create a shared-memory programming environment across the clusters, we can use a virtual shared-memory software layer. Because of the low latency and high bandwidth of the interconnect available within each cluster, there are clear advantages in making the clusters as large as possible. The critical question then becomes whether the latency and bandwidth of the top-level network and the software system are sufficient to support the communication demands generated by different applications. To explore these questions, we have built a virtual shared-memory system using shared-memory multiprocessors and a general purpose network. We implement both single-writer and multiple-writer coherency protocols, adapting those protocols to operate in a clustered environment. We evaluate the overall performance of the system using a variety of applications and examine tradeoffs in clustering and alternative protocols. Our results show that while some well-structured applications can tolerate the latencies of the off-the-shelf interconnects, for a majority of the applications we studied, the high latency of the software system is a serious limitation. For highly clustered configurations, lower bandwidth on a per-processor basis is also a limitation. Thus, increasing the cluster size without scaling the intercluster network leads to poor performance for communication intensive applications. Contrary to the hypotheses of some earlier work, we find that a relaxed single-writer protocol performs better than a relaxed multiple-writer protocol, for four of the five applications we studied. Our results lead us to conclude that it is unlikely that using off-the-shelf networks, virtual shared-memory software, and multiprocessor clusters will lead to a competitive alternative for building large-scale shared-memory multiprocessors that are useful for a wide range of applications.