Cross-Modal Fusing Vision-Language Network for Referring Image Segmentation
Yu Zhao, Wei Wu, Yue Guo · 2023
Reference Image Segmentation (RIS) is a challenging task, swhich aims to match the foreground mask of segmented entities with natural language expressions. In order to solve the RIS problem, we propose two modules: Cross-Modal Fusing Vision-Language (CMFVL) module, Vision and Language Sequence Generation (VLSG) module. VLSG receives visual features and language features, generates feature sequences through attention mechanism, and then sends the sequences to CMFVL. CMFVL module first establishes global attention to vision features through Transformer and then inputs vision features and feature sequences generated by VLSG into Transformer for context information interaction. Experiments on three benchmark datasets confirm the effectiveness of our method.