Semantic Queries with Transformer for Referring Image Segmentation

Yukun Zhai · 2023

Referring image segmentation aims to segment the target region from an image according to the query language description. One of the main challenges behind this fundamental task is to find a qualitative query representation to index the referred object or stuff. In this study, we introduced a query-based framework with Transformer architecture for referring image segmentation, dubbed SQFormer. It treats the sentence and word embeddings as components of two types of semantic queries: (i) mask queries conditioned on the sentence embeddings and (ii) word queries induced from text inputs, to directly attends to the most relevant areas in the image. The semantic queries are input-specific to diverse language expressions while maintaining the prior knowledge of intrinsic image patterns. Concretely, word queries enable flexible and adaptive interactions between vision-language modalities. Mask queries are obligated to generate a set of prototype masks. Then in Prototype Mask Balance (PMB) module, the prototype masks are weighted sum according to the holistic understanding of language expression to get the final mask prediction. Besides, to better fuse linguistic and visual features, we propose a language-aware feature pyramid network (LA-FPN) to enhance the cross-modal alignment. Extensive experiments show our method surpasses the previous state-of-the-art approaches on RefCOCO, RefCOCO+, and G-Ref datasets.

Read the paper · More papers on PaperTik