Processor memory integration: how it affects scalable multiprocessors
Josep Torrellas, Liuxi Yang · 1997
Shared memory multiprocessors provide an easy programming model and the potential to achieve high aggregate performance. To make a shared memory multiprocessor scalable, its memory is distributed among its processing nodes and the nodes are interconnected by a network. The current trend in microprocessor technology is to bring memory closer and closer to the processor. Eventually, each microprocessor will a memory module on-chip. Today''s microprocessors are so fast that the only economically viable choice for designing multiprocessor is to use microprocessors. In this thesis, we argue that this progressive integration of processor and memory has significant implications on the design of DSM systems. To exploit this integration best, we claim that we need to redesign the nodes and reorganize the whole machine. We propose an architecture where memories are configured as caches and directory controllers are moved off-chip. The directory controllers are built out of the same hardware as the computing nodes and, therefore, can be considered non-computing nodes. The function of these non-computing nodes is to support the cache coherence protocol and to backup the application''s data in their memories. Because off-the-shelf processors are so fast, these non-computing nodes can manage the coherence operations and the storage of data in memory in software. A high-level evaluation of the proposed architecture shows that it is significantly better than idealized versions of the traditional COMA and NUMA organizations.