M-segclip: Enhancing SegCLIP with RM-MLP for Open-Vocabulary Semantic Segmentation

Yuhang Zhang, Yanqiu Che, Shanshan Li, Ruofan Wang · 2024

With the emergence of large-scale visual language models represented by CLIP, we have broken the great divide between vision and language. It has become possible to project features from visual models to semantic space consisting of a large amount of text. At the same time, semantic segmentation has opened the door to transferring learned visual knowledge to open vocabulary semantic segmentation. In this paper we improve and optimize a model with open vocabulary segmentation capability based on the CLIP architecture, we design an RM-MLP module and insert this module after the semantic group module in the image encoder, which is used to optimize the semantic regions aggregated by the semantic group module, and then used to generate the final segmentation results. The final experimental results show that the optimized model achieves higher segmentation accuracy on PASCAL VOC 2012 (+1.14% mIoU), and COCO (+0.43% mIoU).

Read the paper · More papers on PaperTik