Investigating Design Trade-Off in S-NUCA Based CMP Systems
Pierfrancesco Foglia, F. Panicucci, Cosimo Antonio Prete, Marco Solinas · CINECA IRIS Institutial research information system (University of Pisa) · 2009
A solution adopted in the past to design high performance multiprocessors systems that were scalable with respect to the number of cpus was the design of Distributed Shared Memory (DSM) multiprocessor with coherent cache, whose coherence was held by a directory-based coherence protocol. Such solution permits to have high level of performance also with high numbers of processors (512 or more). Modern systems are able to put two or more processors on the same die (Chip Multiprocessors, CMP), each with its private caches, while the last level caches can be either private or shared. As these systems are affected by the wire delay problem, NUCA caches have been proposed to hide the effects of such delay in order to increase performance. A CMP system that adopt a NUCA as its shared last level cache has to be able to maintain coherence among the lowest, private levels of the cache hierarchy. As future generation systems are expected to have more then 500 cores per chip, a way to guarantee a high level of scalability is adopting a directory coherence protocol, similar to the ones that characterized DSM systems. Previous works focusing on NUCA-based CMP systems adopt a fixed topology (i.e. physical position of cores and NUCA banks, and the communication infrastructure) for their system and the coherence protocol is either MESI or MOESI, without motivating the reasons of such choices. In this paper, we present an evaluation of an 8-cpu CMP system with two levels of cache, in which the L1s are private of each core, while the L2 is a StaticNUCA shared among all cores. We considered three different system topologies (the first with the eight cpus connected to the NUCA at the same side, the second with half of the cpus on one side and the others at the opposite side, the third with two cpus on each side), and for all the topologies we considered MESI and MOESI. Our preliminary results show that processor topology has more effect on performance and NOC bandwidth occupancy than the coherence protocol.