A reconfigurable liw architecture and its compiler

Rajiv Gupta · 1987

Matching an application to an architecture in structure and size is a way of achieving higher computation speed. Most of the work done in this area focuses on the mapping of applications to fixed architectures. However, no single type of architecture is suitable for all applications. This dissertation presents a combination of a compiler and a reconfigurable long instruction word (RLIW) architecture as an approach to the matching problem. Configurations suitable for execution of different parts of a program are determined by a compiler, and code is generated for both reconfiguring the hardware and performing the computation. Relying on a compiler to detect and schedule parallelism has the advantage that the hardware need not be overly complex and also leads to a reduction in run-time overhead. The RLIW architecture effectively utilizes the fine-grained parallelism detected in programs by compiler techniques. The long word instructions control the operation of processing and memory modules in the system. High memory bandwidth is achieved by dividing the memory into several memory modules that operate in parallel. In order to reduce the data transfer between processing modules and global data memory modules, reconfigurable interconnections among the processing modules are provided which permit direct communication. The compiler uses new techniques including region scheduling, generation of code for reconfiguration of the system and memory allocation techniques to achieve improved performance. The program representation used is an extension to the program dependence graph (EPDG) which divides a program into regions usually consisting of several basic blocks. The region scheduler, guided by the estimates of the parallelism, transforms the EPDG to enable detection and redistribution of parallelism in program regions. Algorithms for packing operations into long word instruction and techniques for effectively assigning memory modules to the operands required by an instruction are developed. Results of the experiments conducted to evaluate the system indicate that speed-ups of 60-300% can be obtained both for scientific and non-scientific programs. The reconfigurable nature of the architecture is responsible for a substantial part of the speed-up. Also results indicate that the major problem of memory bottleneck faced in designing parallel systems is successfully attacked.

Read the paper · More papers on PaperTik