Referring Image Segmentation via Bottleneck Vector-Based Cross-Modal Fusion and Edge-Aware Enhancement
Jiancun Chen, Shulei Zhao, Chenchen Zhou, Binqi Li · 2025
Recently, Referring Image Segmentation (RIS) has drawn significant attention, aiming to segment target objects in images that match the input natural language descriptions. Existing works still face challenges in handling multi-scale objects, cross-modal feature fusion, and fine-grained edge segmentation of targets, which can be attributed to insufficient comprehension and utilization of visual and linguistic information. To address this, we design a Progressive Cross-modal Bottleneck Fusion Network with Local Edge Refinement(PCB-FLR), introducing language guidance during visual encoding and decoding to extract key visual information. Specially, a Cross-Modal Attention Fusion Module Based on Bottleneck Vector(CMAFM-BV) is developed for selective integration of multi-modal features. Finally, a novel Local Edge Refinement Module(LERM) is proposed to refine target boundaries. Extensive comparative experiments demonstrate the superiority of our model, and clear ablation studies validate the effectiveness of our key modules.