Learning Spatiotemporal Features for Video Semantic Segmentation Using 3D Convolutional Neural Networks
Jiamin Chen, Mingchen Wang, Shang Jiang, Bin Huang, Hongbo Sun · 2022
In recent years, significant progress has been made in still image segmentation. However, applying these advanced algorithms to each video frame requires extensive calculation. In this paper, we made two main contributions. The first contribution is a new dataset, we made a human semantic segmentation video dataset based on the Refer-YubeVOS dataset. It provides a benchmark for evaluating video semantic segmentation models. The second contribution is to propose a video semantic segmentation architecture suitable for spatiotemporal feature learning and a method for modifying 2D networks into 3D networks. The trials showed that the 3D network outperforms the 2D network on our dataset. And it is concluded that 3D HRNetV2 has the best performance, with an mIoUvof 61.72%, 14.89% higher than 2D HRNetV2.