Quantization and Acceleration of YOLOv5 Vehicle Detection Based on GPU Chips
Yuxi Yang · 2024
The development of Artificial Intelligence (AI) has raised new requirement for computing chips. The great amount of calculation in deep learning model needs specially designed computing platforms, which is impossible on terminal chips and needs model compression. Model quantization is one way of model compression that can reduce the amount of calculation effectively. This paper is written on the INT8 acceleration of YOLOv5 vehicle detention based on Nvidia Tesla T4 GPU: after YOLOv5 vehicle detention model is trained, one can use TensorRT to accelerate and quantize it. With MAP reducing only 0.6, the overall performance improves by 418%, and the deployable off-line model generated is only 30% of the original size.