Td-VOS: Tracking-Driven Single-Object Video Object Segmentation
Shaopan Xiong, Shengyang Li, Longxuan Kou, Weilong Guo, Zhuang Zhou, Zifei Zhao · 2020
This paper presents an approach to single-object video object segmentation, only using the first-frame bounding box (without mask) to initialize. The proposed method is a tracking-driven single-object video object segmentation, which combines an effective Box2Segmentation module with a general object tracking module. Just initialize the first frame box, the Box2Segmentation module can obtain the segmentation results based on the predicted tracking bounding box. Evaluations on the single-object video object segmentation dataset DAVIS2016 show that the proposed method achieves a competitive performance with a Region Similarity score of 75.4% and a Contour Accuracy score of 73.1%, only under the settings of first-frame bounding box initialization. The proposed method outperforms SiamMask which is the most competitive method for video object segmentation under the same settings, with Region Similarity score by 5.2% and Contour Accuracy score by 7.8%. Compared with the semi-supervised VOS methods without online fine-tuning initialized by a first frame mask, the proposed method also achieves comparable results.