High-speed soft-processor architecture for FPGA overlays

Charles Eric LaForest · TSpace (University of Toronto) · 2015

Field-Programmable Gate Arrays (FPGAs) provide an easier path thanApplication-Specific Integrated Circuits (ASICs) for implementing computingsystems, and generally yield higher performance and lower power than optimizedsoftware running on high-end CPUs. However, designing hardware with FPGAsremains a difficult and time-consuming process, requiring specialized skillsand hours-long CAD processing times. An easier design process abstracts awaythe FPGA via an "overlay architecture", which implements a computing platformupon which we construct the desired system. Soft-processors represent the basecase of overlays, allowing easy software-driven design, but at a large cost inperformance and area. This thesis addresses the performance limitations ofFPGA soft-processors, as building blocks for overlay architectures. We first aim to maximize the usage of FPGA structures by designing Octavo, astrict round-robin multi-threaded soft-processor architecture tailored to theunderlying FPGA and capable of operating at maximal speed. We then scaleOctavo to SIMD and MIMD parallelism by replicating its datapath and connectingOctavo cores in a point-to-point mesh. This scaling creates multi-local logic,which we preserve via logical partitioning to avoid artificial critial pathsintroduced by unnecessary CAD optimizations.We plan ahead for larger Octavo systems by adding support for architecturalextensions, instruction predication, and variable I/O latency. These featuresimprove the efficiency of hardware extensions, eliminate busy-wait loops, andprovide a building block for more efficient code execution.By extracting the flow-control and addressing sub-graphs out of a program'sControl-Data Flow Graph (CDFG), we can execute them in parallel with usefulcomputations using special dedicated hardware units, granting betterperformance than fully unrolled loops without increasing code size.Finally, we benchmark Octavo against the MXP soft vector processor, theNiosII/f scalar soft-processor, and equivalent benchmark implementationswritten in C synthesized with the LegUp High-Level Synthesis (HLS) tool, anddirect Verilog Hardware Description Language (HDL) implementations. Octavo'shigher clock frequency offsets its higher cycle count, performing roughlyon-par with MXP, and within an order of magnitude of the performance of LegUpand Verilog solutions, but with an order of magnitude area penalty.

Read the paper · More papers on PaperTik