CRESO: CLIP-Based Referring Expression Segmentation for Object using Text Prompt

Subin Park, Zhegao Piao, Yeong Hyeon Gu · 2025

Text-based image segmentation is the task of segmenting specific objects in an image based on user-provided text prompts. To improve the performance of existing models, this paper emphasizes accurately embedding text conditions into image features and effectively capturing interactions between objects in complex visual scenes. In this paper, we propose the CRESO (CLIP-Based Referring Expression Segmentation for Objects) model using text prompt which builds upon the state-of-the-art CLIPSeg model. Our proposed CRESO model achieves superior accuracy in image segmentation by leveraging three key methods: the Feature-wise Linear Modulation (FiLM) Layer, Transformer Layer, and multi-scale feature fusion. Specifically, the FiLM Layer adjusts fine-grained image features using text embeddings to enable precise object segmentation based on text conditions, while the Transformer Layer utilizes a self-attention mechanism to learn contextual relationships between objects within an image, enhancing segmentation accuracy for complex and overlapping structures. Furthermore, multi-scale feature fusion enables the segmentation of objects with diverse sizes and intricate structures. Experimental results on the OCID- VLG dataset demonstrate that the CRESO model outperforms competitive baseline across diverse evaluation metrics and precision thresholds. This paper highlights the potential of the CRESO model for effectively addressing applications requiring accurate image segmentation guided by text prompts.

Read the paper · More papers on PaperTik