BSTNet for Content-Fixed Image Harmonization
Yumin Zhi, Youmei Zhang, Bin Li, Dazheng Zhou, Shuhui Yang, Hailong Meng · 2024
Image harmonization aims to enhance the visual coherence of composite images, where composite images refer to placing an object from an image onto another one. Most traditional deep learning-based image harmonization methods utilize encoder-decoders to extract semantic features to achieve foreground-background harmony. With the proposal of RainNet, an increasing number of researchers focus on the foreground-background style feature transfer for image harmonization. While these approaches have demonstrated visually commend-able performance, they neglect that the semantic features learned by convolutional neural networks contain two factors: content and style, which cause the foreground texture features to change. For the above problem, we propose the Content-Fixed Background Style Transfer(CF-BST) module to preserve the content of the foreground while altering its stylistic features to achieve harmony with the background. The CF-BST module includes three steps: first, foreground features are normalized to obtain foreground content features; second, the VGG encoder extracts background style features. Finally, the foreground content and background style features are fused. In addition, we introduce the Enhanced Contextual Fusion (ECF) module to address the limitations of the U-Net in feature extraction and contextual perception. The ECF module consists of three branches that enhance feature representation by aggregating local and global information through multi-scale receptive fields. We conducted experiments on the iHarmony4 dataset and the results show that our proposed network achieves the state-of-the-art performance in image harmonization tasks.