Combining instruction and loop level parallelism for FPGAs
Steven Derrien, Susmita Sur‐Kolay, Sanjay V. Rajopadhye, Rennes-1 Univ., 35 (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National de Recherche en Informatique et en Automatique (INRIA), 35 - Rennes (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National des Sciences Appliquees de Rennes (INSA), 35 (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA) · OpenGrey (Institut de l'Information Scientifique et Technique) · 2000
The conventional method of compiling perfect loops of uniform dependence programs to FPGA based co-processors yields PE arrays where a processor (PE) executes one instance of the loop body per clock cycle. We develop a transformation framework in which the derived PE can be systematically and automatically pipelined through retiming. We use well known transformations, namely skewing and serialization, which enable us to place an arbitrary number of registers at the PE outputs, which are then moved in to the PE's data-path using circuit retimers provided by commercial CAD tools. Our experimental measurements (based on performance estimates after place-and-route) have been very encouraging. For a number of examples we have seen dramatic performance improvements, speed increases of an order of magnitude with relatively little (always less than 50%) area overhead.