Combined instruction and loop parallelism in array synthesis for FPGAs

Steven Derrien, Sanjay V. Rajopadhye, Susmita Sur Kolay · 2001

Compiling perfect, uniform dependence loops to fpga based co-processors normally yields processor pe arrays where a pe executes one instance of the loop body per clock cycle. We develop a transformation framework in which the derived pe can be systematically and automatically pipelined through retiming. We use well known transformations-skewing and serialization, by which an arbitrary number of registers may be placed at the pe outputs. They are then moved into the pe data-path using standard commerecial circuit retimers. Our experiments (based on performance estimates after place-and-route) have been very encouraging. For a number of examples we have seen dramatic performance improvements: speed increases of an order of magnitude with relatively little (always less than 100%) area overhead.

Read the paper · More papers on PaperTik