Techniques for reducing the memory latency for cc-numa multiprocessors
Marius Pirvu, Laxmi Narayan Bhuyan · 2000
Memory latency is one of the major bottlenecks for currently used high performance computers. This problem is more acute in CC-NUMA multiprocessor machines, where remote memory accesses can span several hundreds of cycles. To address this issue we use two techniques. The first one is based on a forwarding engine embedded in the memory controller, which is capable of anticipating processor requests and pushing the desired data, ahead of time, into the processor caches. The prediction mechanism stems from the observation that, in many applications, adjacent memory lines have a similar sharing pattern. The second approach is to effectively reduce the latency of remote memory operations by designing faster interconnect switches. When the network is congested, one of the factors that can affect the performance of a switch is the link arbitration policy. We show that traditional arbitration policies like round-robin or first-come-first-serve do not offer best performance. Therefore, we devise new policies, like shortest message first, priority based and look-ahead, that are better suited for CC-NUMA multiprocessors. Some parallel applications generate only a light network traffic and, in such cases, the most important parameter is the fall-through latency of the switch. To speedup such applications we devise a super-pipelined switch design together with a dual-mode arbitration technique which are able to halve the latency of a state-of-the-art switch like SGI Spider. A thorough performance evaluation should always take into account the effect of hardware complexity on clock cycle. To achieve this goal, we propose a new simulation methodology based on synthesis. We compare and detail three simulation techniques with different accuracy/speed ratios, pinpointing their advantages and disadvantages. Using this new simulation methodology we find that, even though the hardware complexity of the super-pipelined switch slightly degrades the switch clock cycle, the new design is still able to reduce the execution time of parallel applications.