CDGR: Cross-Modal Dual Graph Reasoning for Weakly Supervised Semantic Segmentation

Jia Zhang, Bo Peng, Xi Wu · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Current Convolutional Neural Networks (CNNs) for Weakly Supervised Semantic Segmentation (WSSS) often have difficulties in discovering distinctive feature locations for each category. Therefore, the pseudo-labels generated from the expanded seed regions are typically incomplete and contain a significant amount of noise. Without additional annotations, the numerous erroneous information will potentially propagate in the segmentation network’s training stage. In this work, we propose a Cross-Modal Dual Graph Reasoning (CDGR) framework to leverage both visual and language knowledge effectively. This framework can capture dependencies between the spatial and the semantic spaces, facilitating the discovery of discriminative feature locations. Specifically, we perform cross-modal graph reasoning between the visual and the language modal graphs to enhance global contextual relationships between pixels in the visual feature map. Additionally, we introduce a graph interaction attention network to thoroughly explore implicit relationships between visual and language graphs. We apply the CDGR network to generate more complete pseudo-labels for the classification network and utilize it in the segmentation network to unleash its self-correcting capabilities. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate the effectiveness of CDGR compared to other state-of-the-art peers. Our code is provided at https://github.com/JIA-ZHANG666/CDGR.

Read the paper · More papers on PaperTik