Creating 3D Bounding Box Hypotheses From Deep Network Score-Maps

Lin Guo, Guoliang Fan, Weihua Sheng · 2019

There are two common paradigms for indoor scene understanding, pixel-level labeling and bounding box generation. The two tasks have a complementary nature but are normally achieved separately with different computational flows. We propose a novel method to bridge the two tasks by creating category-specific 3D bounding box hypotheses from score-maps of any deep networks trained on pixel-level semantic labels along with depth data. Those hypotheses can be further used to locate all objects as different non-overlapping bounding boxes by incorporating high-level knowledge, such as common room settings, co-existence or co-exclusiveness etc. We develop an objective function that involves confidence scores and the depth visibility to initialize and optimize multiple hypotheses for each category-specific score map. Experiment results show that our method significantly outperforms direct bounding box generation using pixel-level labeling.

Read the paper · More papers on PaperTik