Boundary Alignment Loss-Guided Cross-Modal Attention Network for Referring Image Segmentation
Lingyun Zhao, Huifang Sun, Longfei Li, Qihang Sun, Wang Nan, Bing Gu · 2024
Referring image segmentation is an advanced computer vision task that aims to accurately segment specific objects in computer vision images using algorithms to understand natural language descriptions. Previous methods rely on explicit human commands to recognize target objects or categories, and lack the ability to actively reason the thin object boundaries, degrading segmentation performance. To solve this problem, this paper proposes a boundary alignment loss-guided cross-modal attention network for referring image segmentation, namely BACNet. Technologically, we firstly introduce a cross-modal attention coupling module (CACM) to realize the bidirectional interaction between language and visual features, effectively exploring the fine-grained information across different modalities. Then, we design a boundary alignment loss (BAL) to refine segmentation boundaries through a differentiable direction vector prediction within an end-to-end framework, significantly enhancing the accuracy of prediction map in aligning with true object boundaries. Specifically, by guiding the Predicted Boundary Descriptors (PDBs) towards the Ground Truth Boundaries (GTBs), BAL dynamically adapts during training, eliminating the need for explicit boundary detection and progressively improving the model performance. Experiments demonstrate the effectiveness of the proposed method on challenging referring image segmentation datasets, exhibiting higher segmentation quality.