Saliency-Guided and Feature Fusion with Transformer for Weakly Supervised Semantic Segmentation
Xiaoman Cai, Wenwu Wang, Lei Zhu · 2024
We propose a saliency-guided Transformer-based weakly supervised semantic segmentation network to address the issues of limited target range and unclear boundaries in the localization of target regions in current weakly supervised semantic segmentation networks. Firstly, by adjusting the weights of attention heads, the model can more completely cover the tar-get area when generating class activation maps. Simultaneously, a saliency-guided self-attention fusion module is introduced to provide more reliable boundary information for the generated class activation maps. To further enhance the performance of weakly supervised semantic segmentation, an online gradient clipping mechanism is designed to effectively filter redundant gradient information by setting appropriate thresholds. Finally, class activation map generation and semantic segmentation ex-periments were conducted on the Pascal VOC dataset, resulting in more complete class activation maps. By obtaining more comprehensive class activation maps, we address the issue of in-sufficient supervision information in weakly supervised semantic segmentation, thereby improving the model's accuracy.