Exploring training recipes and transformer neural networks for optical flow estimation

Prajnan Goswami · 2022

Dense optical flow deals with the motion of all the pixels from one frame to the nextframe in a video sequence. It has become the underlying algorithm for many downstream computer vision tasks such as video frame interpolation, velocity estimation, object track- ing, etc. Until now, a wide variety of optical flow models have been proposed. However, these existing models use different training techniques, data augmentations, and dataset scheduling. As a result, it is difficult to understand if the performance gains comes from the design of the model or due to its training recipe. This thesis aims to provide a comprehensive study of various training recipes on four generations of optical flow models to understand the sources of improvement, and recommend a baseline recipe to train new models and improve existing models. FlowNetC, PWC-Net, RAFT, and a pure Transformer model are the four models considered for the thesis study. However, an end-to-end Transformer model for optical flow is not available at the time of writing the thesis. Hence, we use existing vision transformers with state-of-the-art performance in other dense prediction tasks and compare their efficacy for optical flow estimation.--Author's abstract

Read the paper · More papers on PaperTik