Cooling the Hot Sets: Improved Space Utilization in Large Caches via Dynamic Set Balancing
Mainak Chaudhuri · 2008
Multi-megabyte on-chip last-level caches are commonplace in high-end computing platforms. Even though these caches are often designed to have very high associativity, they suffer from non-uniform utilization of the sets leading to a high volume of conflict misses. Clustering of physical addresses to a few hot sets happens partly due to poor locality in the access stream and partly due to a mismatch in the access pattern and the virtual address to physical address translation algorithm. In this paper, we propose the first fully dynamic mechanism to improve the utilization of sets by adaptively migrating data blocks from hot sets to the relatively cold sets. We present robust and scalable algorithms for identifying the hot sets and suitable cold sets for holding the migrated data blocks that flow from the hot regions. We discuss a number of optimizations on the basic design to address different aspects of such a mechanism, thereby progressively improving the performance. Our detailed execution-driven simulation results show that with just 5.5% extra book-keeping overhead in a 2 MB 16-way set-associative L2 cache, the dynamic block migration mechanism reduces execution time by 12% on average (geometric mean) for nine memory-intensive applications selected from the SPEC 2000 and SPEC 2006 benchmark suites. Further, when applied to an eight-core chip-multiprocessor with a shared 4 MB 16-way set-associative L2 cache, our technique reduces execution time by 18.1% on average (geometric mean) for a set of multi-threaded kernels and applications. We also present a thorough energy analysis of the proposed cache architecture and a quantitative evaluation of how it interacts with an aggressive multi-stream stride prefetcher.