Referring Image Segmentation: An Attention-based Integrative Fusion Approach

Swetha Cadambi, S Shruti, Yerramilli Sri Ram, T. Renuga Devi · 2024

The aim of Referring Image Segmentation (RIS) is to determine an object’s precise boundaries using language descriptions. A significant challenge in RIS lies in integrating visual and linguistic modalities for accurate segmentation, and this paper tackles it by leveraging the integration of multiple fusion structures. A Multimodal Fusion Tree Structure (MFTS) is used to co-relate between the modalities hierarchically, and a Multi-level Fusion Encoder (MFE) module is proposed to model deeper interactions among those multimodal features derived from MFTS and to sieve out the extraneous noise. Further, a Feature Guided Refinement (FGR) module is leveraged to refine multimodal features at different scales and arrive at a final high-quality segmentation mask. The competency of the model is demonstrated by its superior performance over multiple cutting-edge models on three benchmark datasets. The profound impact of the coherent integration of multimodal structures, namely MFTS and MFE, culminates in a striking enhancement in performance by improving the quality of the segmented mask.

Read the paper · More papers on PaperTik