Image Generation with Data Augmentation and Multimodal Conditional Control for Semantic Segmentation

Leilei Wang, Renjie Lu, Jun Yu · 2025

With the rapid advancement of semantic segmentation models, the accuracy of semantic segmentation has significantly improved. However, obtaining a large amount of finely annotated data for training requires substantial resources and time. Image generation, leveraging existing semantic annotations, can produce synthetic images that correspond sufficiently and possess high realism and diversity. In this study, we investigate the data augmentation effects of synthetic images on semantic segmentation models and introduce a pioneering simulation - based approach combining synthetic image generation with traditional data augmentation. Specifically, we apply traditional data augmentation techniques to existing semantic annotations, use the augmented semantic annotations as inputs for image generation, and generate entirely new synthetic images through simulation - driven processes. Through a Multimodal Conditional Control, we ensure the accuracy in the synthetic images, resulting in a usable synthetic dataset. The experimental results reveal substantial enhancements in dataset expansion and accuracy improvement for semantic segmentation models through our simulation - integrated approach. For instance, with Mask2Former on the ADE20K dataset, the accuracy increased from 48.7 % to 51.42%.

Read the paper · More papers on PaperTik