Cascade Cost Volume Multi-View Stereo Network with Transformer and Pseudo 3D
Jiacheng Qu, Shengrong Zhao, Hu Liang, Qingmeng Zhang, Tingshuai Li, Bing Liu · 2023
Learning-based Multi-view Stereo (MVS) and stereo matching methods typically construct 3D cost volumes based on the camera frustum of the reference view. Regularization and regression of the cost volume are performed to obtain a depth map. However, the resolution of the output depth map is limited by the computational cost, and when performing feature extraction, the characteristics of convolution local perception make it impossible to capture global context information. In this paper, we propose CTPMVSNet by using the Global Feature Aware Transformer (GFT) to aggregate global context information within and across images. In order to make better use of GFT, we use Deformable Convolution Module (DCM) to ensure a smooth transition of the extracted feature range. In addition, in the cost volume regularization stage, to improve efficiency and generation accuracy, we design a lightweight regularization network with integrated pseudo-three-dimensional convolution, and our experiments on multiple dataset have achieved promising results.