Vectorized Winograd’s algorithm for Convolution Neural networks

Yuekai Zhao, Jianzhuang Lu, Xiaowen Chen · 2021

Winograd’s algorithm has demonstrated its advantages in accelerating the inference of convolution neural networks. It reduces the number of multiplications in convolution and has achieved success on many accelerators, but there is no practical method to map the Winograd algorithm on a vector processor. With the increasing demand for low-latency inference of convolutional neural networks in data centers, finding effective ways to map the Winograd algorithm for small convolutional kernel networks is essential. In this paper, we propose a mapping method for Winograd’s algorithm on the vector processors. We use scalars to calculate the filter data and obtain a multi-dimensional vector through replication and expansion. Images are calculated as vectors, which improves the program’s parallelism. Then, we build up a three-stage pipeline, including the data read, scalar calculation, and vector calculation. In order to balance the length of the stages, we treat six tiles as a block and calculate them together. The shape of the block can be changed to adapt to different sizes of convolutions. Using output stationary data flow, we temporarily store the intermediate results in the registers to reduce the computational tasks. Finally, we use a double buffering mechanism to reduce the impact of data transmission. Experiments show that the vectorized Winograd algorithm achieves a speedup of 2.2 times compared with direct convolution and the computational efficiency is 44.69%.

Read the paper · More papers on PaperTik