PAI-FCNN

Lansong Diao, Jiang Zhao, Hao Liang, Chang'an Ye, Kai Chen, Li Ding, Shunli Dou, Meng Sun, Lixue Xia, Jiansong Zhang, Wei Lin · 2019

We describe the FPGA subsystem of the Platform of Artificial Intelligence (PAI) in Alibaba Group, called PAI-FCNN. PAI-FCNN plays the role of a heterogeneous back-end for CNN inference, together with other CPU, GPU and ASIC subsystems in PAI. Driven by various business needs, we built PAI-FCNN from scratch since two years ago. We present our experience from FPGA/compiler design and implementation, to system evaluation and deployment. In particular, in order to address three practical challenges: (1) Efficient processing for diverse operators and model structure such as Deconv, Dilated Conv, Up-sampling, PReLu and Concatenation. (2) Serving multiple highly-different models on single FPGA hardware. (3) Competitive performance with alternative GPU or ASIC solutions, we extensively perform joint software & hardware design to optimize system efficiency across multiple CNN models, which includes model reconstruction in compiler software and flexible data access in data-flow CNN processor. We also incorporate reduced precision and model retraining to boost system capacity. Using U-net as an example, on Xilinx KU115 chip, with the help of 74.9% efficiency on Int16-precision hardware (with 3.226TOPS capacity) and 72.9% efficiency on mixed-int8/int3-precision hardware (with 14.746TOPS capacity), we achieve slightly better throughput and 2X higher power efficiency than P4.

Read the paper · More papers on PaperTik