Weakly Supervised Semantic Segmentation via Contextual Cross-Attention

Zhan Xiong, Fons J. Verbeek · 2024

Weakly supervised semantic segmentation (WSSS) using image-level labels faces two challenges: Class Activation Maps (CAMs) lack explicit background modeling and lose spatial information during pooling. However, post-processing for background inference often results in imprecise separation. Moreover, the absence of geometric regularization in standard loss functions hinders accurate boundary delineation. To tackle these limitations, we propose a novel end-to-end WSSS framework integrating a conditional probabilistic model with a geometric prior. Our framework is initiated by explicitly modeling the background using Bernoulli distributions trained with contrastive self-supervised learning. We then introduce a variational total variation (TV) term to refine background regions and improve the boundary fidelity. Foreground class probabilities are subsequently modeled with Categorical distributions conditioned on the refined background. This conditional foreground modeling is implemented as a novel Contextual Cross-Attention (CCA) module, which can capture long-range dependencies among foreground regions. Our hierarchical framework offers insights into WSSS representation learning and significantly improves segmentation accuracy on the PASCAL VOC dataset.

Read the paper · More papers on PaperTik