HoloDreamer: Holistic 3D Panoramic Scene Generation From Text Descriptions

Haiyang Zhou, Xinhua Cheng, Wangbo Yu, Yonghong Tian, Li Yuan · IEEE Transactions on Visualization and Computer Graphics · 2025

3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, creating 3D scenes using only text prompts has become viable, thereby significantly advancing research in text-driven 3D scene generation. Prevailing methods typically employ the diffusion model to generate an initial local image, followed by iteratively outpainting the local image to gradually generate scenes. Nevertheless, these outpainting-based approaches are prone to producing globally inconsistent results with low completeness, restricting their broader applications. To tackle these problems, we introduce HoloDreamer, a framework that begins by generating a high-definition panorama to holistically initialize the full scene, and leverages 3D Gaussian Splatting (3D-GS) for rapid 3D scene reconstruction, thereby facilitating the creation of view-consistent and fully enclosed 3D scenes. Specifically, we propose Stylized Equirectangular Panorama Generation, a pipeline that combines multiple diffusion models to enable stylized and detailed equirectangular panorama generation from complex text prompts. Subsequently, Enhanced Two-Stage Panorama Reconstruction is introduced, conducting a two-stage optimization of 3D-GS to inpaint the missing region and enhance the integrity of the scene. Comprehensive experiments demonstrated that our method outperforms prior works in terms of overall visual consistency and harmony, as well as reconstruction quality and rendering robustness when generating fully enclosed scenes.

Read the paper · More papers on PaperTik