Improving memory hierarchy performance by store-load renaming and locality-guided parallelization

Peng Lü, Jih-Kwon Peir · 2005

Memory hierarchy performance has been a dominant factor for processors' performance. The problem has been exacerbating because of the increasing performance gap between processors and memory hierarchy. Data communications among producer and consumer instructions in a program can go through fast speed registers with minimum delay. However, the number of registers is usually limited to save all intermediate data. In the case of not enough registers, a program writes the data into memory hierarchy and reads the data out whenever it is ready. The data communications through memory hierarchy incur extra delays that further degrade processor performance. Our observation is that we can represent these data communications in instruction syntax correlation format instead of memory address. Specifically, store-load and load-load syntax correlations can be represented in the form of base register ID plus the displacement value. Using this decoding information and taking advantage of spatial locality among memory references, we proposed a new memory layer, signature buffer, with novel addressing memory mechanism to avoid long delay of address calculation and cache memory accessing. Performance evaluations based on an Alpha 21264-like pipeline using SPEC2000 benchmarks show that an IPC (Instruction-Per-Cycle) improvement of 12–17% is possible using a small 8-entry signature buffer. The memory behavior of programs is complicated when they run on processors with multithreads sharing one cache. First, the memory reference sequence among multiple threads can be either constructive or disruptive. The memory behavior is constructive when multiple threads share the data that were brought into the shared cache by one or another. In other words, memory reference locality exhibited in the original single thread may be transformed and/or further enhanced among multiple threads. The memory behavior is disruptive when multiple threads have rather distinctive working sets and compete with the limited memory hierarchy resources. Second, the memory usage or footprint may increase among multiple threads based on the fact that certain data structures or local variables may be replicated to improve parallelism. The increase of memory footprint alters the reference behavior and demands larger caches to hold the working set. In this dissertation, we demonstrate a methodology for memory behavior analysis for emerging multimedia, artificial intelligence and database applications running on multithreaded processors and show their memory behavior results.

Read the paper · More papers on PaperTik