A distributed predictive cache for high performance computer systems

Thomas M. Alexander, Gershon Kedem · 1995

The execution time of programs that have large working sets is substantially increased by the overhead of retrieving data from the memory system. Even when a first level cache is integrated with the CPU, the memory overhead may increase the total execution time by 50-300%. In this dissertation we present a cost effective memory system that uses a novel address predictor to reduce the latency and a distributed second level cache to increase the bandwidth of the memory system. We have studied the inter-arrival times of requests to main memory as well as the memory address request patterns for a workload of large scale programs. From this study we observed that the CPU address request patterns to main memory repeat during program execution. We propose a 1st order Markov model to abstract the CPU memory address request behavior. The model is used to build a dynamic hardware based address predictor, i.e., 1st order Memory Address Predictor (MAP), to predict the next CPU request. The 1st order MAP is comprised of a Prediction Table that stores an approximation of the state transition matrix of the model, a learning and self-correcting unit that refines the contents of the Prediction Table by tracking the CPU behavior and a Prefetch Unit that uses the Prediction Table to predict CPU requests. The 1st order MAP is designed to predict regular memory access patterns such as those generated by matrix programs as well as much more complicated address request patterns like those generated by linked-list based programs. To facilitate the transfer of large blocks of data from the slow main DRAM memory to the fast second level (L2) cache, the L2 cache is integrated into the DRAM array to form a distributed cache. The resulting large bandwidth between the DRAM array and the L2 cache is utilized by the 1st order MAP to move large blocks of data between the DRAM array and the L2 cache, concurrent with CPU execution. We present two practical architectures, i.e., Cached DRAM & MAP and Enhanced DRAM & MAP, using our methodology to build high performance memory systems. Both systems are designed around commercially available memory ICs. Accurate event-driven simulation of the memory systems show that the Enhanced DRAM & MAP scheme has an average memory overhead of 8%, which is a 71% reduction when compared to a conventional 256K byte L2 cache. In addition, the Enhanced DRAM & MAP needs a total of 160k bytes of SRAM which is 38% less than the 256K bytes needed for the conventional L2 caches.

Read the paper · More papers on PaperTik