VLC-UNIT: Unsupervised Image-to-Image Translation with Vision-Language Classification

Yuying Liang, Huakun Huang, Bin Wang, Lingjun Zhao, Jiantao Xu, Chen Zhang · 2024

In recent years, unsupervised image-to-image translation technology has garnered widespread attention for its creative expression, accompanied by enhanced data augmentation capabilities. However, current methods are heavily influenced by dataset distributions, resulting in reduced performance on data-scarce samples and inconsistencies in fine-grained categories and target styles, leading to the generation of pseudo categories. To tackle these challenges, we introduce a novel framework called Vision-Language Classification (VLC-UNIT) for UNsupervised Image-to-Image Translation. By leveraging prompts from large vision-language models, VLC-UNIT enhances the semantic understanding of images and improves the consistency and accuracy of image translations. Experimental results demonstrate our VLC-UNIT outperforms existing state-of-the-art techniques.

Read the paper · More papers on PaperTik