CrossTR-Raft :Dense optical flow estimation based on cross attention mechanisms

Zimeng Liu, Xingming Wu, Liu Zhong, Jing Zhang, Haosong Yue, Weihai Chen · 2024

Opticalflow essentially represents the displacement of a certain pixel between two consecutive image frames. It has become one of the most significant research topics in the field of computer vision, enabling the understanding of motion and scene dynamics. Recently, a lot of optical flow estimation frameworks have emerged and had a notable impact such as Recurrent All-pairs Field Transforms (RAFT). RAFT achieves notable performance in optical flow estimation by extracting per-pixel features through convolutional neural networks, constructing multi-scale 4D correlation volumes, and iteratively updating the flow field using a recurrent unit. Moreover, Transformers have become popular due to their strong representational capacity. However, Transformers commonly contain a self-attention rather than cross attention mechanism. This paper introduces RAFT of Cross-attention Transformer, i.e., CrossTR-RAFT, a novel approach to optical flow prediction, which incorporates a cross attention mechanism into the RAFT backbone framework. The cross attention is proficient in capturing long-range dependencies between different spatial locations, thus facilitating the learning of more robust and accurate flow representations. Our proposed approach can obtain an End-Point Error (EPE) of 0.65 and 1.08 on Sintel clean and final train splits respectively, which means a 5.80% and 10.74% reduction on the corresponding dataset compared to that without cross attention. The experimental results on benchmark datasets demonstrate that our proposed approach achieves state-of-the-art performance in terms of accuracy and robustness, showcasing its capability to handle challenging scenes and intricate motion patterns.

Read the paper · More papers on PaperTik