EM-LAST: Effective Multidimensional Latent Space Transport for an Unpaired Image-to-Image Translation With an Energy-Based Model
Giwoong Han, Jinhong Min, Sung Won Han · IEEE Access · 2022
For an unpaired image-to-image translation task to work well, the latent space of each image domain must be well designed, and the codes of the style must be translated toward the target with the preservation of the parts corresponding to the source content. In general, most variational autoencoder (VAE)-based models use one-dimensional latent space. However, to apply high dimensional methodologies such as vector quantization, which is a recent trend in deep learning, it is necessary to be able to control multidimensional latent space. Among the VAE-based models using relatively complex multidimensional latent spaces, we apply the methodology of an energy-based model (EBM) with vector quantized VAE v2 (VQ-VAE-2) as the main model. Our study shows that among the latent spaces that represent each image domain, the importance of each feature at the top and bottom latent spaces corresponding to the high- and low-level features must be interpreted differently for proper translation. We also present various analyses and visual outcomes of multidimensional latent space transport.