Neural Radiance Fields for Sparse Satellite Images Leveraging Semantic and Geometry Consistency
Wenbo Sun, Zhi Yang Gao, Yanzhang Li, Jun Zhu · IEEE Transactions on Geoscience and Remote Sensing · 2025
How to achieve accurate scene reconstruction models using limited views has long been a key research topic in the fields of photogrammetry and computer vision. Under sparse view conditions, the lack of sufficient semantic and geometric priors often results in suboptimal performance of reconstruction algorithms. Recently, neural radiance fields (NeRF) have gained a lot of attention in satellite scene reconstruction. However, existing NeRFs for satellite scenes require many views to render novel views and generate DSMs, which conflicts with the scarcity of satellite imagery. The key to improving performance degradation lies in how to extract rich scene semantics and geometry from a limited number of views, thereby enhancing the understanding of the scene. This paper presents a novel NeRF-based method that leverages semantic and geometry consistency to enhance the NeRF performance with sparse satellite images. Specifically, we introduce a cross-view semantic consistency loss to enhance the semantic understanding capability of the NeRF model with limited views. We utilize a vision language foundation model for remote sensing, RemoteCLIP, to extract semantic embeddings and ensure the consistency of semantic features across different views of the same scene. Additionally, we propose a geometric consistency loss to ensure the geometric consistency between surface normals and scene depth, thereby ensuring the scene geometric accuracy when NeRF model generates images from unseen views. Extensive experiments on various urban satellite scenes demonstrate that our proposed method achieves superior performance on both novel view synthesis and DSM generation tasks under sparse-view training conditions.