Effective communication methods for many core architectures with on chip networks in the absence of cache coherence

Pablo Reble, Rudolf Mathar, Matthias Müller · RWTH Publications (RWTH Aachen) · 2016

If the trend of integrating more and more cores to a single die continues, general-purpose processors with thousands of cores may be expected in the near future. The result of this development would be many-core architectures that will inevitably create new challenges for the scalability of common synchronization and communication methods. It is commonly believed that a reevaluation of established concepts is needed to address such research challenges. Because the importance of efficient communication for processors with many cores cannot be underestimated, the focus of this dissertation is on the analysis of efficient communication methods to exploit locality, and on low-latency of architectures that follow the network-on-chip paradigm and provide software-controlled memory with a low latency and a high throughput. As a member of the Many-core Applications Research Community, the RWTH Aachen University started, in 2010 the projects MetalSVM and iRCCE to explore new communication concepts and system-software support for future many-core systems. These projects – the results of which are presented in this dissertation – include the development of software concepts for communication and synchronization in the absence of hardware-cache coherence. Intel’s Single-chip Cloud Computer (SCC) represents the first x86-based many-core processor. It implements a new communication concept with alternative support for on-chip consistency control and explicit communication. These attributes, in combination with a flexible and fine-grained memory control, have enabled experiments and the development and verification of low-level communication concepts today, and it thereby guides the development of future systems. To further explore the scalability of low-level software for this kind of architecture, this dissertation includes consideration of the design of a transparent virtual extension. A major achievement of res-ulting low-level communication framework is a full working prototype of a cluster of clusters-on-a-chip, which can emulate a many-core processor with more than two hundred cores. This prototype has enabled a deeper analysis of new many-core communication concepts, and has uncovered potential for optimization. Major performance improvements could be achieved by the combination and further development of well-known mechanism and software techniques.Moreover, a communication model is essential if we are to analytically explore the limits of many- core processors that follow the network-on-chip paradigm and implement a low-memory abstraction by providing software-controlled on-chip memory. The SCC has been used to support analysis of methods and to verify the accuracy of our concepts when used to model communication. The experimental hardware shares basic characteristics with attributes of future many-core architectures that result, for example, in a combination of stacked memory and a tight integration of the fabric interconnect. Such similarities create opportunities for the applicability of the effective communication methods to future architectures. Fundamental requirements for the efficient communication methods that are developed, evaluated and parametrized in this work include configurable caches for direct on-chip memory access or at least finer-grained cache control for data movement. Moreover, a low-level contention model is de-veloped to evaluate different synchronization concepts and to derive optimizations for many-cores with remote direct memory access.

Read the paper · More papers on PaperTik