End-to-end Compilation is All FPGAs Need: A Unified Overlay-based FPGA Compiler for Deep Learning

Kai Qian, Haodong Lu, Yinqiu Liu, Zexu Zhang, Kun Wang · 2025

Field-Programmable Gate Array (FPGA) has shown great application potential in deploying Neural Networks (NNs) due to the characteristics of programmability, low power consumption, etc. However, deploying NNs on FPGA is non-trivial because (1) Mainstream NNs pose significant FPGA architecture design challenges due to their large number of parameters, complex operations, and the need for data optimization, and (2) Supporting the deployment of different machine learning frameworks to FPGA requires significant manual effort, consuming a large amount of time. In this paper, we propose AutoCompiler, a unified compiler for mapping NNs to different FPGAs, along with overlay techniques to enable fast and efficient implementation. To the best of our knowledge, we are the first work to support both Deep Neural Networks (DNNs) and Transformer-based networks for overlay-based FPGA deployment. AutoCompiler comprises three integrated enablers: (1) Model Translator, built on top of a topology-based NNs representation, which can optimize the topology and data representation of the models from an algorithmic level based on different hardware configurations, e.g., DSP utilization, (2) Instruction Generator, which generates pipeline data streams according to various FPGA resource configurations by manipulating the instruction set at the upper level rapidly, and (3) End-to-end optimization, which moves as much of the computational processes as possible onto the FPGA chip and minimizes the interaction between CPU and FPGA. Extensive experiments on various Xilinx FPGAs show that AutoCompiler outperforms state-of-the-art overlay-based compiler by 1.2× - 1.35× and same-level GPUs by 1.15× - 1.59× for classic DNN models, and ViT inference, respectively.

Read the paper · More papers on PaperTik