DVS-3D: Diffusion-based novel view synthesis and 3D object reconstruction from a single image
Xu Kui, Tiejun Wang, Xiaoran Guo, Xiaoyan Hu, Lingmei Tao, Chaoyang Wu · Journal of Computational Design and Engineering · 2025
Abstract Current three-dimensional (3D) reconstruction methods typically rely on multi-view image datasets or costly 3D scanning devices. While these approaches can produce accurate 3D structures, they face practical challenges, including data acquisition difficulties, slow processing speeds, and high computational demands. Moreover, traditional single-view methods struggle to infer occluded regions and reconstruct high-frequency geometric details in complex scenes, often producing 3D models with compromised shape accuracy and view-dependent texture inconsistencies. To address these limitations, we propose Diffusion-based View Synthesis for 3D reconstruction (DVS-3D), a single-image 3D reconstruction framework based on Stable Video Diffusion. Our framework integrates CLIP-embedding, Variational Graph Autoencoder, and a Multi-Adversarial Visual Transformer within a U-Net architecture. This design effectively captures both global and local image features, allowing explicit adjustment of camera poses to generate multiple novel views of the target object. Compared to traditional fixed-viewpoint methods, DVS-3D achieves more precise and flexible 3D representations, overcoming challenges associated with limited viewpoint information and complex geometric reasoning. Extensive experiments on Objaverse and OmniObject3D datasets, along with additional evaluations on the google scanned objects dataset, demonstrate that our model outperforms state-of-the-art methods such as Zero123, Zero123-XL, and StableZero123 in handling complex textures and occlusions. Our approach consistently achieves significant improvements across various evaluation metrics, validating its effectiveness in generating high-fidelity 3D models from a single image.