FINN-T: Compiling Custom Dataflow Accelerators for Quantized Transformers
Christoph Berganski, Felix Paul Jentzsch, Marco Platzner, Max Kuhmichel, Heiner Giefers · 2024
Transformers evolved to dominate the state of the art in almost any deep leaning domain, and usually involve very large models that are computationally expensive for training and inference. Smaller Transformer variants are still useful, particularly for a wide range of embedded applications. Existing compilers for FPGAs often lack support for core operations of the Transformer, such as the attention mechanism. In this work, we leverage the FINN framework to enable the automatic synthesis of custom-tailored FPGA accelerators for quantized Transformer models. We describe the hardware design of the scaled dot-product attention operation following the streaming dataflow paradigm and the integration into the compiler infrastructure of the FINN framework. We demonstrate a small-scale design space characterization exploring scaling behavior and resource utilization of our implementation, followed by two exemplary case studies, covering a radio signal classification use case as well as text generation with small GPTs, to demonstrate the end to end toolflow starting from quantization aware training up to deployment on the device. We evaluate our approach in terms of resource utilization, throughput, latency, accuracy and perplexity.