MT-CLIP: Open-Vocabulary Segmentation Based on Dual Mask Promotion with Text Embedding
Lina Han, Juan Wei, Tianping Li, Jianlong Zhang · 2025
Recent open-vocabulary semantic segmentation[1] methods depend on mask generators to extract features aligned with textual embeddings; however, the resulting masks often lack sufficient classification accuracy. The present paper is concerned with the presentation of MT-CLIP, a semantic segmentation framework that utilises dual mask prompts and text embeddings and which is notable for its openness with regard to vocabulary. Through this approach, the visual and textual representations of CLIP are synergistically optimized, leading to improved alignment in the visual-textual embedding space. We achieved performance improvements in mIoU on the open-vocabulary segmentation datasets A-847, A-150, and PAS-20.