Blocking Linear Algebra Codes for Memory Hierarchies

Steve Carr, Ken Kennedy · 1989

. Because computation speed and memory size are both increasing, the latency of memory, in basic machine cycles, is also increasing. As a result, recent compiler research has focused on reducing the effective latency by restructuring programs to take more advantage of high-speed intermediate memory (or cache, as it is usually called). The problem is that many real-world programs are non-trivial to restructure, and current methods will often fail. In this paper, we present some encouraging preliminary results of a project to determine how much restructuring is possible with automatic techniques. 1. Introduction. Over the past decade we have seen dramatic reductions in the cycle times of microprocessors, while memories for the same processors have been growing in size. These two trends have yielded computer systems in which memory latency is quite large in terms of basic machine cycles---latencies of 10 to 20 cycles are not unusual. To address this problem, system designers have incorpor...

Read the paper · More papers on PaperTik