WSSS-DM: Advancing Weakly Supervised Semantic Segmentation by Harnessing Diffusion Models
Chuanguo Shen, Zhaofeng Niu, Bowen Wang, Guangshun Li, Liangzhi Li · 2024
The development of the highly observed field of semantic segmentation is hindered by time-consuming and complex manual mask annotations. This study seeks to address this issue by introducing a new method, WSSS-DM, which allows end-to-end weakly supervised semantic segmentation based solely on text descriptions. Using the cross attention word-pixel scores derived from the denoising self-network of the diffusion model, masks are generated. Different masks collected after screening are then combined onto the same image, forming an image-label pairs corresponding to the original image. Simultaneously, we categorize based on the quality of labels, differentiating the degree to which explainable features of images can be accurately obtained. In the experiments performed on the COCO captions dataset, the high-quality category yielded an mIoU of 37.4 and an mAcc of 42.17, thereby substantiating the effectiveness of our method.