Multicore cache hierarchies: design and programmability issues
Ramón Doallo, Óscar Plata · Concurrency and Computation Practice and Experience · 2013
Multicore cache hierarchies: design and programmability issuesWelcome to this special issue of the journal Concurrency and Computation: Practice and Experience on Multicore Cache Hierachies -Design and Programmability Issues, which contains three original manuscripts.Caches have been playing an essential role in the performance of single-core systems due to the gap between processor speed and main memory latency.First level caches are strongly restricted by their access time but current processors are able to hide most of their latency using outof-order execution as well as miss overlapping techniques.On the other hand, last levels of the cache memory hierarchy are not so dependable on their access time but on their locality issues.The locality in lower levels is filtered by the upper levels.As requests going down in the memory hierarchy, they require a greater number of cycles to be satisfied, so it becomes more difficult to hide the latency of last-level caches.In multicore systems, their importance is even larger due to the growing number of cores that share the bandwidth that this memory can provide.In an attempt to make a more efficient usage of their caches, the memory hierarchies of many chip multiprocessors present last-level caches, which can be allocated across threads and part of them may be private to a thread while other parts may be shared by multiple threads.Then, caching techniques will continue their evolution during next years in order to tackle the new challenges imposed by multicore platforms and workloads.A clear indicator of the current interest of the research community in new techniques for optimizing the performance and power consumption of multicore cache hierarchies is the organization during last years of specific sessions devoted to these topics at top international conferences on computer architecture and parallel computing.This special issue contributes to this promising field with extended and carefully reviewed versions of selected papers first from the International Workshop on Multicore Cache Hierachies -Design and Programmability Issues, which was held as part of the 10th IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA 2012) in Madrid (Spain), and second from all the community in the field as the Call for Papers was also open to contributions that were not sent to the mentioned workshop.The first contribution, by O. G. Lorenzo et al. [1], presents a set of three hardware counter (HC)-based tools to characterize memory access of parallel codes in symmetric multiprocessors.This toolkit simplifies accessing and programming HCs, which are included in modern microprocessors.Hardware counters are used to obtain information about memory accesses in a parallel code at very low cost.This information is presented to the user in a friendly way.The first tool can be used to automatically monitor the memory accesses of a system and to analyze a code even if the source is not available.The second tool allows the user to insert in a source code, in a simple and transparent way, the instructions needed to monitor and manage HCs, so specific parts of the code can be analyzed.The third tool takes the information gathered by the aforementioned tools, processes it and displays it graphically, allowing the user to adjust the level of detail.The aim of these tools is to characterize the memory accesses of parallel codes in multicore systems, in which the cache hierarchy can greatly influence the performance.Cache coherence techniques have been introduced to enable fast access while preserving the data coherence but these coherence protocols are critical in hard real-time systems.Because the frequent inter-cache communication leads to unpredictable interferences between the cores, the system's timing behavior is hard to analyze.Pyka et al. [2] propose a new hard real-time capable strategy for multicore systems called on-demand coherent cache (ODC 2 ).The technique is based on marginal hardware extensions compared to non-coherent caches and the use of common synchronization techniques.ODC 2 provides coherent accesses to cached shared data as well as caching of private data.Because the presented strategy does not induce interferences between local caches, ODC 2 is capable for hard real-time systems.