A Semi-Supervised Video Object Segmentation Method Based on ConvNext and Unet
Dan Han, Yuelei Xiao, Pengyu Zhan, Tao Li, Mengyu Fan · 2022 41st Chinese Control Conference (CCC) · 2022
Semi-supervised video object segmentation task, i.e., separating objects in a video from the background given the mask of the first frame during segmentation. Based on the One-shot Video Object Segmentation (OSVOS) method, this paper proposes a semi-supervised video object segmentation method based on ConvNext and UNet. In this method, the original VGG was first replaced with pure ConvNet models called ConvNeXt, which outperformed VGG in feature extraction and improved by 13.4% on the ImageNet dataset; then, using the UNet architecture, it Compared with FCN, a different feature extraction method is adopted, which makes it more suitable for large image segmentation and medical segmentation; finally, Focal Loss is used to dealing with the unbalanced binary classification problem, which solves the problem that the iteration of simple samples is slow and cannot achieve the best results, and balances the uneven proportion of positive and negative examples, making our method more competitive. Our experimental results on the DAVIS2016 dataset show that the feature extraction effect of ConvNeXt is better than that of VGG in this dataset, and the UNet architecture is more suitable for video object segmentation tasks than FCN. Focal Loss also improves the experiment significantly. A J&F of 91.7 was achieved, enhancing the original OSVOS by 80.2%. Experiments show that this series of modifications are very effective and substantially improves the technical level.