Open-vocabulary 3D Semantic Segmentation With 3D Region Mask Proposal and 2D-3D Visual Feature Ensemble
HyeLim Bae, Incheol Kim · Journal of Institute of Control Robotics and Systems · 2024
Open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. In this study, we propose a novel open-vocabulary 3D semantic segmentation model, OV-3DRENet, to address the limitations of existing models. Unlike existing 3D semantic segmentation models that perform point-level categorization, the proposed model performs region-level categorization using Mask3D as a 3D region mask proposal module to generate multiple class-agnostic point-cloud regions. The proposed model uses OpenScene, a pre-trained open-vocabulary point-cloud segmentation model, as a 3D point encoder to extract language-aligned 3D visual features for each region from the scene point cloud. Furthermore, it adopts OpenSeg, a pre-trained open-vocabulary image segmentation model, as a 2D pixel encoder to extract language-aligned 2D visual features for each region from multiview scene images. Finally, our model applies a novel 2D 3D visual feature ensemble scheme to allocate well-matched open-vocabulary class labels to point-cloud regions. By conducting various quantitative and qualitative experiments using a large benchmark dataset, ScanNet v2, we demonstrate the superiority of the proposed model.