Object detection using part based semantic segmentation
Răzvan Bogdan ITU, Radu Gabriel Danescu · 2021
Monocular vision systems are increasingly popular in driving assistance applications as they are easy to set up and do not require precise calibration or synchronization. The downside of monocular vision is the lack of 3D information, which makes the task of identifying individual objects that are close together in the image space difficult. The lack of 3D information must be compensated by high accuracy classification of the image data. This paper proposes a novel way of detecting objects using fully convolutional neural networks followed by lightweight geometric based post processing. The fully convolutional neural network has four semantic segmentation outputs corresponding to quarters of individual objects. Therefore, each pixel of the input image will be classified as either belonging to a top left, a top right, a bottom left, or a bottom right region of a whole object. If the object is occluded and only a few of the four regions are visible, the component pixels will still be labeled correctly. Based on the multiple outputs of the neural network, the pixels are grouped into connected regions using a clustering algorithm aware of the relations between the object’s quarters. The accuracy of individual obstacle instances is similar to the accuracy of the results obtained from instance segmentation networks, while the demand of resources and the number of trainable parameters is significantly reduced.