Cooperative system software and architectural mechanisms for efficient distributed shared memory multiprocessing

John B. Carter, Chen-Chi Kuo · 1999

The performance of scalable hardware distributed shared memory (DSM) multiprocessors is often limited by the amount of time spent handling remote memory accesses. Thus, the value of a DSM architecture is directly related to the extent to which observable remote memory latency can be reduced to an acceptable level. This dissertation describes the design of a hybrid DSM architecture, called Adaptive Simple COMA (AS-COMA), that has integrated two traditional scalable shared memory architectures: cache coherent nonuniform memory access (CC-NUMA) and simple cache-only memory architecture (S-COMA). We have developed three novel techniques in AS-COMA to dramatically reduce the amount of kernel overhead and the number of remote misses caused by needless thrashing of the page cache: (i) an LRU with backoff replacement algorithm, (ii) an S-COMA-preferred initial allocation policy, and (iii) an algorithm for adapting to oscillating memory pressures. AS-COMA's LRU with backoff replacement algorithm can identify when hot pages start to evict each other from the page cache and dynamically reduce the rate of page remappings to avoid further thrashing. The S-COMA-preferred initial allocation policy can improve the performance of hybrid architectures by accelerating their convergence to pure S-COMA behavior when the memory pressure is low. The adaptive algorithm for oscillating memory pressures detects relieving memory pressures and lowers relocation thresholds to allow more relocations. As a result, AS-COMA can exploit the best features of S-COMA at low memory pressure and CC-NUMA at high memory pressure. To demonstrate that AS-COMA is a cost-effective architecture for the next generation of DSM multiprocessors, we have carefully compared AS-COMA with contemporary design alternatives. AS-COMA outperforms CC-NUMA and S-COMA under almost all conditions, and outperforms other hybrid architectures by up to 17% at low memory pressure and up to 90% at high memory pressure.

Read the paper · More papers on PaperTik