Learning Video Correspondence using Appearance Module for Target Tracking
Jin-Mo Choi, Jeany Son, Sangjoon Park · 2021
We introduce a new method for self-supervised video correspondence matching to effectively track targets in battlefield situations. Specifically, we propose Appearance module that enforces to maintain a target appearance between two consecutive frames in a local window, where an affinity matrix is computed on high-resolution feature maps with a small search window. It has an advantage of mitigating an over-fitting problem of a conventional affinity matrix that only predicts motions by reducing ambiguity of labels in a pretest task. Also the proposed method preserves the detailed shape of an object by handling high resolution information while high computational costs due to the correlation filter is alleviated by a small search window. Our experimental results on the DAVIS2017 dataset showed the significant performance improvement over the baseline.