Implementing VTA, a tensor accelerator on Flow-in-Cloud
Kazuei Hironaka, Kensuke Iizuka, Hideharu Amano · 2021
A multi-FPGA system Flow-in-Cloud (FiC) consists of nodes with a mid-class cost-efficient FPGA and Raspberry Pi3B connected with high speed serial links. It aims to implement a large scale AI applications which is difficult to be implemented on a single FPGA by dividing the target into a number of boards. On the other hand, overlay domain specific architecture receives attention for implementing complicated application programs easily on the FPGA. Here, we focus on an open-source AI compiler framework Apache TVM, and its implementation for FPGA, VTA(Versatile Tensor Accelerator). In order to use FiC through the TVM, we implemented it on FiC and executed the ResNet-18 inference benchmark. The preliminary evaluation results showed that VTA on FiC-SW achieved up to 10 times performance compared to the execution on ARM Cortex-A54 software.