MVoxTi-DNeRF: Explicit Multi-Scale Voxel Interpolation and Temporal Encoding Network for Efficient Dynamic Neural Radiance Field
Weiqing Yan, Yanshun Chen, Wujie Zhou, Runmin Cong · IEEE Transactions on Automation Science and Engineering · 2024
Neural radiance fields have revolutionized the field of novel view synthesis, achieving remarkable results. However, traditional approaches based on implicit representations, particularly those built upon NeRF, suffer from slow rendering speeds due to the need for numerous MLP evaluations. Recently, there has been a promising shift towards explicit representations using voxel grids, which has significantly improved reconstruction times for static scenes. Nonetheless, extending these methods from static scenes to dynamic scenes is a non-trivial task as it requires accounting for the changing geometry and appearance of the scene over time. In this paper, we propose an efficient dynamic neural radiance field with multi-scale explicit voxel interpolation and temporal encoding. We leverage an explicit voxel structure to store the 3D dynamic features, while employing a lightweight MLP to estimate the displacement, thereby significantly enhancing the reconstruction speed. In our canonical module, we incorporated temporal information encoding in density estimation and color estimation to rectify the error estimation of the displacement in deformation module. In addition, a multi-scale voxel interpolation is designed to accommodate large-scale motions while meticulously capturing intricate details in small-scale motions in the density estimation module. In experiment, we evaluate our MVoxTi-DNeRF method on both synthetic and real scenes, where it achieves superior or comparable rendering quality compared to state of the art methods, while remaining computationally efficient (more than$60\times $faster than the original DNeRF). More experiment results and test code are available athttps://github.com/CHenYYff/MVoxTi-DNeRFNote to Practitioners—Neural Radiance Fields(NeRF) can create highly detailed and realistic 3D models from 2D images via a neural network. It is particularly popular for its capacity to generate novel views of a scene, enabling it to synthesize images from previously unobserved viewpoints. NeRF has found applications in a wide range of fields, including virtual reality, augmented reality, gaming, and so on. In this study, we introduces an efficient dynamic NeRF method. First, our network leverages an optimized explicit voxel grid to store 3D dynamic features and employs a lightweight MLP to decode these deformation features, significantly accelerating the training process. Second, to correct the error estimation related to deformation displacement, we introduce encoding of temporal information into density and color estimation in our canonical module, which fortifies the canonical field’s perception of temporal information. Third, we utilize multi-scale voxel interpolation to capture different-scale motion in the density estimation module, where minor motions are modeled using nearby voxels, while motion within a broader range is captured through more distant voxels. The consideration of multi-scale voxel features diminishes the detrimental impact caused by inaccurate displacement estimation.