Joint 2D Object Detection and 3D Reconstruction via Adversarial Fusion Mesh R-CNN

Zihan Zhou, Qinghan Lai, Shuai Ding, Song Liu · 2021

Joint 2D object detection and 3D reconstruction is an essential computer vision task to get more accurate detection and representation model of the target object. We proposed a novel joint 2D object detection and 3D reconstruction model that enhances the ability of the 2D object detection and the 3D reconstruction, called Adversarial Fusion Mesh Region Convolutional Neural Networks (AFM R-CNN). Our proposed model introduces the Deep Convolutional Generative Adversarial Network (DCGAN) to generate adversarial images and input the real and adversarial images into the object detection module GA-RPN to determine the position and anchor box of the target object. Next, to make better use of the two-dimensional information of the image, the voxel conversion and Fusion model Pix2Vox is introduced to fuse the two types of image features and generate coarse voxels. Afterwards, to differentiate the voxel information more efficiently, we use the Principal Neighborhood Aggregation network (PNA) model in 3D model refinement. The contrast experimental results on the open domain dataset (Pix3D) with baseline models demonstrate the effectiveness of AFM R- CNN in joint 2D object detection and 3D reconstruction task.

Read the paper · More papers on PaperTik