PGDM: Multimodal Panoramic Image Generation with Diffusion Models
Depei Liu, Hongjie Fan, Junfei Liu · 2024
Panoramic image generation enhances visual representation and facilitates immersive experiences across various fields. However, Existing models readily gives rise to error accumulation and loop closure. To overcome these limitations, we introduce PGDM, a diffusion-based framework that facilitates the concurrent generation of visually cohesive panoramic images. In order to achieve modality alignment, PGDM adopts the vision-language model as a Multimodal Module for extracting visually representative text-aligned features. Additionally, we design the Multi-View Attention with Camera Pose mechanism between each UNet layer to enhance the coherency of panoramic images. Our extensive experimental results demonstrate that PGDM achieves the state-of-the-art performance while simultaneously maintaining diversity and consistency across multiple view images.