Hardware-Aware Quantization and Performance Evaluation for Tensor Accelerator VTA
Huazheng Zhao, Shanmin Pang, Yinghai Zhao, Haihong Lang, Yunxin He, Hengxiang Xu, Fan Gao, Kuizhi Mei · 2023
In order to improve the quantization inference accuracy of the deep learning compiler framework TVM(Tensor Virtual Machine), we proposed a quantization deployment inference scheme, which is designed for accelerated INT8 calculation of deep neural network based on VTA, using quantization-compilation collaborative optimization algorithm. Firstly we analyzed the limitations of quantization scale calculation in the original Post-Training Quantization scheme of TVM. According to the Hardware-Aware strategy, we carried out the fixed-point calculation optimization of scale and other model operators in the process of QAT (Quantization Aware Training) to optimize the integer tensor acceleration characteristics of VTA(Versatile Tensor Accelerator), meanwhile the fixed-point inference error including quantization error was considered. The scheme is hardware-friendly, reduces the data memory occupancy, and can be used for multiple tasks such as image classification and object detection. Experiment results show that compared with the original TVM scheme, the model inference speed and accuracy are significantly improved, which provides an efficient and lightweight solution for deep neural network model deployment using fixed-point accelerator.