Photographic image synthesis with improved U-net

Gang Liu, Jundong Si, Yanzhong Hu, Shan Li · 2018

Photographic image synthesis is the process of creating images that are accurate representations of a real scene with some form of image description and is a new application field o f deep l earning technology. The U-net is a frequently used convolutional network architecture in deep learning technology and is mainly used for image segmentation. Currently, the U-net has been used for photographic image synthesis. However, the attenuation of image description information, the low memory capacity of the network and the checkerboard artifacts in the output images cause the low quality of the images synthesized by the U-net. In order to solve the above problems, this paper presents an improved U-net (IUN) for photographic image synthesis. In IUN, the image description map is resized to the multiple resolutions and the resized images are merged into the different layers of IUN for enhancing the image description information. The resize-convolution layers in IUN replaces the deconvolution layers in the U-net to reduce the checkerboard artifacts. The proposed learnable residual blocks increase the memory capacity of the network. IUN can be easy trained end-to-end with the image description map and the corresponding photographic image. In experiments on datasets, two types of image description information, semantic and sketch, are used to synthesize the photographic images. Compared to the optimization-based networks, cascaded refinement networks (CRNs) and other methods, IUN gives more realistic results. Additionally, our network achieves a balance between memory capacity and time cost.

Read the paper · More papers on PaperTik