Evaluation of a Stall Cache: An Efficient Restricted On-chip

Klaus Erik Schauser, Krste Asanović, David A. Patterson · 1991

In this report we compare the cost and performance of a new kind of restricted instruction cache architecture the stall cache against several other conventional cache architectures. The stall cache minimizes the size of an on- chip instruction cache be caching only those instructions whose instruction fetch phase collides with the memory access phase of a preceding load or store instruction. Many existing machines provide a single cycle external cache memory [6, 17, 2]. Our results show that, under this assumption, the stall cache always outperforms an equivalent sized on-chip instruction cache, reducing external memory access stalls by approximately 10%, In addition we present results for a system using an on-chip data cache, and for one with a double width data bus and short instruction prefetch buffer. On most RISC architectures, the only instructions that access memory are loads and stores. Usually variables and indeterminate results can be kept in registers, reducing the number of data accesses as compared to memory-memory or register-memory architectures. However, studies have shown that loads and stores still account for around 25-40% of all instructions executed [9]. If a processor is to avoid memory access stalls, the memory subsystem must be capable of delivering at least one instruction word every cycle while servicing data accesses. A CPU that has only a single bus to external cache memory might encounter a structural hazard and have to stall during the memory access phase of a load or store instruction. There are several ways to reduce the number of these load and store stalls. One solution is to have separate data and instruction busses; another possibility is to have caches on chip. In this report we evaluate the cost and performance of a number of existing solutions for removing these stalls against a new proposal: the stall cache [7]. We first present the architectural alternatives, then develop a cost and performance model for each, and finally present experimental results for traces taken from programs in the SPEC benchmark suite.

Read the paper · More papers on PaperTik