Referring Image Segmentation with Two-Stage Multi-Modal Interaction

Zhenhua Wang, Linwei Ye · 2024

The objective of referring image segmentation is to extract referred entities from an image using a particular natural language sentence. The main idea for this task is interacting textual and visual features to build multi-modal relationships. The prior state-of-the-art methods mainly focus on local multilevel intermediate feature interaction or global text-to-image alignment, which might result in insufficient interaction for capturing global multi-modal information exchange or finegrained referred object details, respectively. To overcome this issue, we introduce a referring image segmentation framework with two-stage multi-modal interaction. Specifically, we devise an innovative multi-level cross-modal fusion module to effectively facilitate the interaction of intermediate features of linguistic and visual modalities for fine-grained details of referred objects. Besides, we further align the linguistic and visual information by introducing an elaborate global alignment module for accurately localizing the entire referred objects. The comprehensive experiments conducted on three referring image segmentation datasets illustrate that our proposed two-stage multi-modal interaction framework exhibits a marked superiority over the contemporary state-of-the-art approaches.

Read the paper · More papers on PaperTik