VCNN: A compiler of CNNs based on MLIR for multi-core vector accelerators

Xiaorong Chen, Cheng Li, Zhong Liu B · 2024

Convolutional Neural Network (CNN) is one of the representative algorithms of machine learning and deep learning. Multi-core vector accelerators are becoming increasingly popular due to their high performance and low power consumption. In the previous methods of deploying CNNs to vector accelerators, three issues are of concern. (1) Multi-core tasks are divided according to the image data dimension, which has limitations in low-latency real-time detection application scenarios. (2) Many memory optimization strategies only consider the impact of tensor size or tensor life, ignoring the inherent computational characteristics of operator types. (3) Manual mapping methods require a lot of engineering effort and are prone to introducing errors. Therefore, the VCNN proposed in this paper is a compiler based on MLIR for vector accelerators. The VCNN can automatically map CNN models to vector accelerators. It includes (1) a Multi-core Parallel Convolution algorithm (MPC) that is more suitable for low-latency real-time detection application scenarios. (2) An Adaptive Memory Reuse method (AMR). It not only considers the size or lifetime of tensors but also the inherent computational characteristics of operator types. (3) A quantization pass that quantizes the model and reshuffles the data. Experimental results show that the VCNN can effectively utilize the parallel processing capabilities of hardware when performing convolutions of different sizes, and achieve a parallel computing efficiency of up to 95.75%. In addition, we evaluated the performance of AlexNet, VGG16, and Yolov5s models. The results show that VCNN has a computational efficiency of more than 2X+ that of TVM, TensorRT, and OnnxRuntime, and a power efficiency of more than 7X+ that of them. Additionally, the VCNN inference model using Float16 is twice as energy efficient as that using Float32. Furthermore, the VCNN achieves a higher memory reuse rate compared to traditional memory reuse methods.

Read the paper · More papers on PaperTik