CNN-based Object Detection on Low Precision Hardware: Racing Car Case Study
Nicolo De Rita, Alessandro Aimar, Tobi Delbrück · 2019
Increasing interest in deep learning and convolutional neural networks resulted in the last years in multiple techniques aiming to improve their accuracy, training speed, and inference speed. At the same time, their computational cost triggered the design of several dedicated hardware architectures, aiming to handle the elevated number of operations neural networks require with minimal power budget, often exploiting reduced precision arithmetic. In this case study, we analyzed how several techniques can be merged together in the design of a track detector for a self-driving racing car, illustrating a step-by-step procedure required to adapt several theoretical works to a real-world scenario. Compared with the best previous detector, the new Proteins cone detector is optimized for low-precision deep learning accelerators. It runs 50% faster on GPU than the previous detector and at a simulated 272.5 FPS on a 1 W ASIC or at 16.4 FPS on a 12 W FPGA and achieves a detection score 16% higher than the previous implementation.