MEWS: Semantic image segmentation with multiclass extreme weak supervision

A. Apostolidis, Vasileios Mygdalis, Matthaios Dimitrios Tzimas, I. Pitas · Neurocomputing · 2026

Unsupervised image segmentation methods typically assume zero a-priori knowledge about the data semantics. This assumption does hold in many practical scenarios where, although the training data might not be annotated, the target semantic image region classes are known. In these settings, text-driven prompting methods for semantic image segmentation offer noticeable improvement in segmentation accuracy over purely unsupervised approaches. However, such approaches are still limited by: a) inherent text-prompt semantic ambiguity, b) ineffective adaptation to target domain distributions, and c) excessive computational and architectural complexity. To address these shortcomings, we propose the Multiclass Extreme Weak Supervision (MEWS) framework for semantic image segmentation. MEWS assumes the availability of extremely few class-based pixel-level image annotations, e.g., few annotated image pixels per class in very few training images. Such pixel-based image prompts are thereby employed to form image region class prototypes. They can be used to leverage low-complexity unsupervised image segmentation architectures to be trained by our novel prototype-based triplet loss that learns discriminative image features by promoting intra-class image feature compactness while enforcing inter-class feature vector separation. Consequently, the proposed MEWS image segmentation architecture leads to increased weakly supervised training efficiency, bridging the performance gap between supervised and unsupervised image segmentation methods. Our experimental results indicate that the proposed methods compare favorably against text-based prompting image segmentation methods. It yields superior image segmentation accuracy in publicly available image segmentation datasets (e.g., Cityscapes), as well as in Natural Disaster Management (NDM) ones. • A novel Multiclass Extreme Weakly Supervised (MEWS) semantic segmentation framework is proposed that generalizes the original binary EWS DNN architecture by utilizing only sparse, per-class, few-pixel per class labelling. • A class prototype-based triplet loss function is designed that pulls same-class prototype feature vectors together, while pushing mean prototype feature vectors belonging to different classes apart. • A multiclass dynamic thresholding mechanism improves contrastive learning, without additional supervision or manual hyperparameter tuning. • MEWS segmentation excels in accuracy, scaling with annotation, ablation study on loss, validated on NDM Sardinia Wildfire dataset.

Read the paper · More papers on PaperTik