Weakly supervised zero-shot cross-modal semantic segmentation

Hsiao-Cheng Lin, Jou-Yi Su, Jing-Ming Guo, Yi‐Chong Zeng · 2025

Weakly supervised semantic segmentation reduces annotation costs by using less detailed data, such as image-level labels or bounding boxes, but often suffers from lower accuracy due to insufficient annotations, leading to classification and boundary errors. Additionally, training weakly-supervised models requires complex algorithms, increasing computational resources and training time. This paper introduces an algorithm that combines a generalization unsupervised segmentation model with zero-shot learning and cross-modal understanding between images and texts. This approach reduces training time and computational costs while improving object boundary recognition. Tested on the PASCAL VOC 2012 dataset, the algorithm achieves a mean Intersection over Union (MIoU) of 77.3% on the test set, with a minor increase in computational speed by only 0.04 seconds per frame.

Read the paper · More papers on PaperTik