Sparse Deep Neural Network Acceleration on HBM-Enabled FPGA Platform

Abhishek Kumar Jain, Sharan Kumar, Aashish Tripathi, Dinesh Gaitonde · 2021

Domain-specific architectures (DSAs) for deep neural networks (DNNs) are becoming mainstream because of high energy-efficiency and performance. Currently, most of these DSAs perform dense linear algebra efficiently by exploiting temporal and spatial locality, high data reuse and regular memory access patterns. However, state-of-the-art DNN models require DSAs with very high amount of compute, interconnect and memory resources. Since sparsity in DNN parameters can be exploited to reduce the resources and complexity of the DSAs, emerging DSAs are adding native support for sparsity [1], [2]. This paper presents a sparse DNN inference engine on HBM-enabled Alveo U280 platform, built around FPGA-optimized DSA blocks. The DSA blocks have native support for completely unstructured sparsity. We demonstrate the effectiveness of proposed inference engine by running inference models of the Sparse DNN Challenge [3]. Results for the challenge benchmarks show that the proposed inference engine and multi-DSA parallelization on HBM-enabled Alveo U280 achieve up to 3.7 billion edges per second inference throughput.

Read the paper · More papers on PaperTik