Monocular 3D Object Detection using Disjoint Generalized Intersection-over-Union Loss

Andrea Zinelli, Luigi Musto · 2020

Three-dimensional object estimation from monocular imagery is a difficult task due to the geometric information loss induced by the perspective projection. Currently, most methods try to deal with this limitation by using deep learning models augmented with some form of 3D reasoning, which often leads to complicated inference pipelines. In this work, we argue that explicit 3D reasoning is not mandatory for good monocular 3D detection performance. To this end, we extend a 2D object detection framework with a small subnetwork responsible for 3D bounding box estimation. To train this module, we introduce a loss function based on the Generalized Intersection-over Union in which each degree of freedom is optimized separately. The resulting approach is simple, modular and straightforward to integrate into existing monocular 2D detection frameworks. Experiments on the challenging KITTI dataset show that our method achieves state-of-the-art performance on the 3D detection task.

Read the paper · More papers on PaperTik