Tolerating First Level Memory Access Latency in High-Performance Systems.
William Y. Chen, Scott A. Mahlke, Wen‐mei Hwu · 1992
Abstract-- In order to improve performance, future parallel systems will continue to increase the processing power of each node in a system. As node processors, though, can execute more instructions concurrently, they become more sensitive to the rst level memory access latency. This paper presents a set of hardware and software techniques, collectively referred toasregister preloading, to effectively tolerate long rst level memory access latency. The techniques include speculative execution, loop unrolling, dynamic memory disambiguation, and strip-mining. Results show that register preloading provides excellent tolerance to rst level memory access latency up to 16 cycles for an issue 4node processor.