Optimized Yolov3 Deployment on Jetson TX2 With Pruning and Quantization

Zhuoxuan Shi · 2021 IEEE 3rd International Conference on Frontiers Technology of Information and Computer (ICFTIC) · 2021

Pruning and Quantization are commonly techniques deployed on deep convolutional networks to accelerate the model and reducing size, taking the resolution as a sacrifice. In embedded systems where computation power is limited due to power and cost constraints, pruning and quantization is mandatory for such networks especially where real-time processing is required, such as image classification in live video streams. In this paper, a pruned and quantized YOLOv3 model is deployed on Nvidia's industrial standard model Jetson TX2, which demonstrated an increase of 5 FPS in image classification via an 640×480 USB camera, while allocating only 16.7% storage on disk.

Read the paper · More papers on PaperTik