Quality-driven model-based design of multi-processor accelerators:an application to LDPC decoders
Jan, Y (Yahya) · TU/e Research Portal · 2012
The recent spectacular progress in nano-electronic technology has enabled the implementation of very complex multi-processor systems on single chips (MPSoCs). However in parallel, new highly demanding complex embedded applications are emerging, in fields like communication and networking, multimedia, medical instrumentation, monitoring and control, military, etc., which impose stringent and continuously increasing functional and parametric demands. The high demands of these applications cannot be satisfied by systems implemented using general purpose processors (GPP). For these applications increasingly complex and highly optimized application-specific MPSoCs are required to perform real-time computations to extremely tight schedules, when satisfying high demands regarding the energy, area, cost and development efficiency. High-quality MPSoCs for these applications can only be constructed through adequate usage of efficient application-specific system architectures exploiting more adequate concepts of computation, storage and communication, as well as, usage of efficient design methods and electronic design automation (EDA) tools for synthesizing the actual high-quality hardware platforms implementing the architectures. Some of the representative examples of these highly-demanding applications include the based-band processing in wired/wireless communication (e.g. the upcoming 4G wireless systems), different kinds of encoding/decoding in communication, image processing and multimedia, 3D graphics, ultra-high-definition television (UHDTV), and encryption applications, etc. These applications require to perform complex computations with a very high throughput, while at the same time demanding low energy and low cost. The decoders of the low density parity check (LDPC) codes, adopted as an advance error-correcting scheme in the newest wired/wireless communication standards, like IEEE 802.11n, 802.16e/m, 802.15.3c, 802.3an, etc., for applications as digital TV broadcasting, mm-wave WPAN, etc., can serve as a representative example of such applications. These standards, for instance, the IEEE 802.15.3c specifies as high as 5~6 Gbps throughput for the upcoming wireless communication systems. For the realization of the so high throughput as several Gbps massively parallel hardware multi-processors are indispensable. These modern complex applications involve massive parallelism of various kinds (e.g. task, data and functional, etc) and complex interrelationships among the data and computing operations. Therefore, an adequate accelerator design for such applications requires a careful exploration and exploitation of various kinds of parallelism and resolution of complex interrelationships between the data and computing operations. The accelerator design for such kind of applications has to involve adequately combined micro- and macro-architecture design for the processors, and the corresponding adequate memory and communication architectures design. Since the processor’s micro-/macro-architecture and the memory and communication architectures are strongly interrelated and cannot be designed in separation, complex mutual tradeoffs have to be resolved among the processor parallelism at the two levels, and the corresponding memory and communication architectures, as well as, among the performance, area and power consumption. For the design of hardware accelerators high-level-synthesis (HLS) methods and tools are often used. However, the HLS methods and tools only support the micro-architecture synthesis of a single processing unit, while not taking into account the macro-architecture, memory and communication synthesis and not accounting for the relationships and tradeoffs among these design aspects, what is necessary in the design of hardware accelerators for highly-demanding applications. To address the issues highlighted above and to resolve the mutual tradeoffs effectively and efficiently, a novel quality-driven model-based hardware multi-processor design method is proposed in this thesis and related design space exploration (DSE) tool that jointly consider the processor, memory and communication architectures, and the possible mutual tradeoffs among them. For the high-throughput requirements that demand massively parallel hardware multi-processor architectures, the communication and memory have usually a dominating influence on all the most important architecture quality aspects, such as delay, area and power consumption. Although some research results on the memory and communication architectures can be found in the literature, these results are for programmable on-chip multi-processor systems that utilize the time-shared communication resources, such as shared buses or Network on Chip (NoC), which are not adequate to sustain the multi-Gbit bandwidth required for the high-end massively parallel hardware multi-processors. Therefore, the research reported in this thesis was especially focused on the memory and communication issues, and proposed several promising generic scalable communication and memory architectures to satisfy the required data transfer bandwidth of the high-end hardware accelerators. In its scope several generic hierarchical partitioned communication and memory architectures were proposed, as well as, their application method to the memory and communication design of massively parallel accelerators that ensure the scalability when applied to massively parallel hardware multi-processors. The proposed novel design method makes it possible to perform an effective and efficient exploration and exploitation of the various tradeoffs between the processing parallelism at the micro- and macro-architecture level, and the corresponding memory and communication architectures, as well as, among the area, performance and power consumption, to arrive at high-quality accelerator architectures. Several novel scheduling, processing parallelism exploration, and the memory and communication architecture exploration strategies are incorporated into the proposed architecture DSE framework. To analyze and evaluate the proposed design methodology and its related DSE framework, a series of extensive case studies are performed through implementing and applying the methodology for several industrial strength applications of the LDPC decoding for the latest communication system standards. These case studies involved extensive architecture synthesis experiments with the LDPC decoder designs for IEEE 802.15.3c LDPC codes. In particular, the results of our experiments clearly demonstrate that neither the fully serial nor the fully parallel micro-architectures are adequate to satisfy the ultra-high performance requirements. The extreme fully serial or fully parallel micro-architectures are also not appropriate from the viewpoint of area and power consumption. To satisfy the ultra-high or ultra-low performance requirements, the combined micro-/macro-architecture exploration is necessary which explores and exploits various partially parallel architecture combinations. The results of the experiments performed confirmed that without considering the micro- and macro-architecture design, as well as, processor, communication and memory architecture design in combination, it is very difficult to arrive at an adequate high-quality multi-processor accelerator. They confirmed that our design methodology adequately supports the design of complex multi-processor accelerators, while taking into account the numerous complex tradeoffs. To our knowledge, despite a more than a decade of research on the hardware accelerators for the highly-demanding applications no similar holistic quality-driven design approach has been proposed. In this thesis, we take into account all the design components jointly as a single design task, as well as, consider the mutual tradeoffs among them and among different design objectives. Finally, using our method, it is possible to implement various high-quality multi-processor accelerators for the highly-demanding applications (e.g. LDPC decoders of practical importance for the newest upcoming wireless communication standards) with a limited human designer effort and in a short time.