S2I-Bird: Sound-to-Image Generation of Bird Species using Generative Adversarial Networks
Jooyong Shim, Joongheon Kim, Jong‐Kook Kim · 2021
Generating images from sound is a challenging task. This paper proposes a novel deep learning model that generates bird images from their corresponding sound information. Our proposed model includes a sound encoder in order to extract suitable feature representations from audio recordings, and then it generates bird images that corresponds to its calls using conditional generative adversarial networks (cGANs) with auxiliary classifiers. We demonstrate that our model produces better image generation results which outperforms other state-of-the-art methods in a similar context.