FPGA-Based Accelerator for Deep Convolutional Neural Networks for the SPARK Environment
Raghid Morcel, Mazen Ezzeddine, Haitham H. Akkary · 2016
Deep Convolutional Neural Networks are known for their high performance. However, their complexity is one of their most challenging aspects. In this paper, we propose a design of an FPGA-based accelerator for the distributed training of convolutional neural networks. The accelerator is intended for use in the SPARK run time environment to accelerate the training of deep convolutional neural networks in the data center. Our accelerator is very energy efficient and achieves 40 to 250 times speedup in the computation of the multi-layer convolution operation, a key operation in the training/inference of deep networks.