Text-Guided Refinement for Referring Image Segmentation
Shuang Qiu, Shiyin Zhang, Tao Ruan · Applied Sciences · 2025
Referring image segmentation aims to segment an object described by a natural language expression from an image. Existing methods perform multi-modal fusion during encoding, typically integrating image and text features before predicting masks via upsampling networks. However, this approach often lacks sufficient multi-modal interactions during decoding, leading to challenges in achieving precise edge predictions for objects with varying scales. Additionally, the isolated interaction between linguistic and visual features at different scales fails to utilize the continuous guidance of language to multi-scale visual features. To address this issue, we propose the Text-Guided Refinement Network (TGRN). It employs a cascaded pyramid structure with a text-guided gating mechanism to enable selective and efficient integration of multi-modal features across multiple scales at the decoding stage. The proposed TGRN offers the following advantages: (a) It enhances information flow across feature scales, improving the network’s capacity to represent multi-scale semantics and achieve accurate segmentation. (b) It leverages text information to guide feature fusion, allowing for strengthened multi-modal interactions and refined edge perception during decoding. (c) It facilitates effective multi-modal information integration through a language-embedded visual encoder. Extensive experiments on three benchmark datasets validate the effectiveness of the proposed approach, demonstrating its superior performance in referring segmentation.