Multi-View Classification and 3D Bounding Box Regression Networks

Christopher Pramerdorfer, Martin Kampel, Mark Van Loock · 2018

We present a method for jointly classifying objects in depth maps and regressing amodal (extending beyond occluded parts) 3D bounding boxes in a way that is highly robust to occlusions. Our method is based on a novel multi-view convolutional neural network architecture with shared layers for both tasks, improving efficiency. The network processes views that encode object geometry and occlusion information and outputs class scores and bounding box coordinates in world coordinates, requiring no post-processing steps. We demonstrate the effectiveness of our method by example of fall detection, presenting a new dataset of 40k samples rendered from 3D models. On this dataset, our method achieves an average classification accuracy above 97% and a regression error below 10 cm at occlusion ratios of up to 90%. The dataset and trained models are publicly available.

Read the paper · More papers on PaperTik