A Self-Attention Network for Stereo Matching
Menglong Yang, Hanyong Wang, Ren Yang · 2024
Learning contextual information has been proved to be conducive to reducing mismatches in ill-posed regions, and many data-driven stereo matching algorithms achieve state-of-the-art performances. Global contextual information is difficult to learn by a shallow convolutional neural network, while a large network brings huge computational cost. This paper proposes a self-attention framework of stereo matching to learn both global and local contextual information. From a rough estimation, the disparity can be extremely refined by applying the proposed attention module. We present two self-attention modes to better learn global and local context synchronously. Experimental results show the proposed algorithm can well predict thin structures, large occluded or textureless regions, and achieves the comparable performance to state-of-the-art methods on the public stereo benchmark with a real-time speed.