Prompts Libra: Enhanced Image Outpainting Diffusion Model With Balanced Bimodal Guidance

Zongyan Zhang, C. L. Philip Chen, Zepeng Su, Tong Zhang · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Image outpainting, a challenging generative task, has advanced significantly with the introduction of text-to-image diffusion models (DM). Despite these advances, DM-based methods frequently encounter the phenomenon in which one modal takes precedence over another, causing the image to be over-guided. Current research relies on manual hyperparameters to achieve bimodal balance. To reduce reliance, Prompt Libra is proposed to automatically balance bimodal prompts during inference and enhance extrapolated images. Given the variation of bimodal cross-attention during DM denoising, we create an adaptive bimodal attention module via attention maps. Furthermore, we design a classifier-free guidance computation based on masked images to improve the semantic control of the masked part and enhance the quality of images. Finally, we propose a semantic transformer to address the problem of quality degradation caused by incomplete prompts. It extracts limited semantics from the source images, which is suitable for scenarios lacking text prompts. Experimental results demonstrate that our method generates images that achieve the state-of-the-art effect on several image quality evaluation metrics while maintaining the image and text prompts in balance.

Read the paper · More papers on PaperTik