Extended data cache prefetching using a reference prediction table

Donglok Kim, Yongmin Kim · 1997

The rapid advances in the VLSI technology and computer architecture have made a significant improvement in the CPU performance possible in the last 30 years. However, the growth rate of the memory system performance has remained slower than that of microprocessors, resulting in the number of CPU clock cycles required to access main memory doubling approximately every 6.2 years. Cache prefetching is one of the most useful techniques to overcome the memory latency problems. Many sophisticated methods such as SPT (Stride Prediction Table) and RPT (Reference Prediction Table) have been shown to predict the prefetch address using the stride (difference in the addresses of the current and previous memory accesses) information extracted from the instruction stream. While these prefetch mechanisms seem to work well for the loops that have a large number of iterations or have enough time for the prefetched data to become available before it is actually used in the next iteration, we have found that the hit-wait cycles (the CPU stall cycles when the load/store instruction hits the cache line that is currently being prefetched) can occupy a significant portion of the cache service time. This means that we can reduce the cache service cycles if the prefetched data arrive earlier. Early arrival of prefetched data can be accelerated in two ways: advancing the perfetch issue timing and increasing the prefetch throughput. In order to address these issues, we present two extensions to the existing RPT (Reference Prediction Table) data cache prefetching. The first algorithm, SSE-RPTP (Small Stride Extended RPT Prefetch), prefetches the next cache line even when the detected stride is smaller than the cache line size, reducing the hit-wait cycles. The second algorithm, XRPTP (eXtended RPT Prefetch), tries to issue a maximum of two cache line prefetches at a time, further reducing the hit-wait cycles and improving the prefetching throughput. The simulation results on 11 SPEC92 programs show that the improvement in the cache service time is 34.3% for the SSE-RPTP and 37.9% for the XRPTP, respectively.

Read the paper · More papers on PaperTik