Generative Multiplane Image (GMPI): Text to Volumetric Representation

Dewei Hu, Dae Yeol Lee, Guan‐Ming Su · 2024

In this paper, we propose a novel framework for generating volumetric representation from a text prompt. Our method leverages Stable Diffusion to initially generate an image corresponding to the text. Then, we construct a multiplane image (MPI) representation of the image, consisting of fronto-parallel planes at discrete depth intervals. Each MPI plane holds RGB textures and alpha opacity (A) information. These planes can be warped and rendered to simulate views on other camera poses. The key challenge of MPI lies in the accurate reconstruction of occluded RGB textures. To address this issue, we first generate occlusion mask of each layer using scene opacity information, then employ an inpainting network to restore textures of these regions under multi-perspective, text-guided supervision. More specifically, we use the inpainted MPI to render novel views on randomly sampled camera poses. Then, we compute score distillation sampling (SDS) loss to update the inpainting network for the rendered scenes' realism and their adherence to the text prompt. Experimental results demonstrate that our method effectively generates MPI with disoccluded textures, providing an immersive volumetric viewing experience for the given text prompt across wide range of viewpoints.

Read the paper · More papers on PaperTik