An Improved Method for Zero-Shot Semantic Segmentation

Kong Kuok Yong, Tan Fong Ang, Chin Soon Ku, Firdaus Sahran, Lip Yee Por · IEEE Access · 2025

Zero-shot semantic segmentation continues to face challenges in effectively handling unseen object classes, despite its critical applications in medical imaging, autonomous driving, and aerial imagery. This paper introduces a hybrid approach that combines three key components: a Swin Transformer backbone, a Flash Masked Attention Transformer Decoder, and a Deformable Pixel Decoder, to enhance zero-shot semantic segmentation accuracy. The Swin Transformer backbone extracts hierarchical image features, the Flash Masked Attention Transformer Decoder refines feature extraction by focusing cross-attention on foreground regions, and the Deformable Pixel Decoder specializes in capturing small object features. Extensive evaluations on benchmark datasets demonstrate the method’s performance, achieving scores of 31.1 mIOU and 71.8 pAcc on ADE20K, 9.3 mIOU and 54.4 pAcc on ADE20K-Full, and 94.7 mIOU and 97.5 pAcc on Pascal VOC, significantly outperforming existing approaches. These results establish the effectiveness of the proposed method in advancing zero-shot semantic segmentation, particularly in scenarios involving diverse and previously unseen object classes.

Read the paper · More papers on PaperTik