Leveraging Monocular Depth and Feature Fusion for Generalized Stereo Matching
Jun Jiang, Xiaoyan Liao, Fan Yang, Ka Chun Cheung, Xin’an Wang, Yong Jun Zhao · 2025
Depth estimation tasks play a crucial role in computer vision, with monocular depth estimation and stereo matching being the main approaches. Following the introduction of end-to-end deep learning methods, stereo matching networks have made significant progress in both accuracy and speed. However, these methods often focus on optimization for synthetic datasets, which results in weak generalization across different target domains, limiting their application in real-world scenarios. Monocular depth estimation networks, trained on a wide range of large-scale datasets, exhibit strong cross-domain generalization capabilities, enabling them to adapt to various scenes, objects, and lighting conditions. In contrast, stereo networks rely on disparity information, which limits their generalization ability and makes them more susceptible to changes in viewpoint and matching accuracy. Therefore, this paper proposes a new method that combines the strengths of monocular and stereo networks to improve the cross-domain generalization ability of stereo networks. To this end, we introduce a new cost volume strategy for fusing depth information and RGB features, and employ cosine groupwise cost volume to enhance the robustness of the matching process. Experimental results demonstrate that the proposed fusion method improves the network's performance in cross-domain tasks, validating the potential of monocular and stereo feature fusion in enhancing the network's domain generalization ability.