Cache Design Options for a Clustered Multithreaded Architecture
Rajeev Garg, Ali A. El‐Moursy, Sandhya Dwarkadas, David H. Albonesi, Jude A. Rivers, Viji Srinivasan · UR Research (University of Rochester) · 2005
The design of the memory hierarchy in a multi-core architecture is a critical component since it must meet the capacity (in terms of bandwidth and low latency) and coordination requirements of multiple threads of control. Most previous designs have assumed either a shared L1 data cache (e.g., simultaneous multithreaded architectures) or L1 caches that are private to each individual processor (e.g., chip multiprocessors (CMPs)) with coherence maintained across the L1s at the L2 level. A shared L1 cache has the benefit of potentially increasing cache capacity for threads with non-uniform working sets but the disadvantage of higher access latency from remote clusters/cores and the potential for conflicts among threads. Private caches have the benefit of lower L1 access latency but the disadvantages of reduced effective cache size and of coherence overhead. In this paper, we focus on the design of the L1 cache as being a critical design component especially for multithreaded/parallel workloads. We identify the factors contributing to reduced performance, and examine ways in which each factor---locality of access, capacity, as well as migratory, multiple-reader, and read-write sharing patterns---can be addressed in a low-overhead fashion. We demonstrate that direct access to remote L1 cache banks provides the flexibility to address thread-specific capacity issues as well as to provide improved sharing support for active read-write sharing. We propose a novel selectively replicating (SR) adaptive protocol to handle locality and capacity issues in addition to recognizing and adapting to migratory, multiple-reader, and read-write sharing access patterns. The SR cache design is able to meet or beat the performance of the coherent cache, improving performance by up to 44% relative to a coherent cache (with an average of 5% across all applications) when using 8 threads.