Exploiting Temporal Information in Real-time Portrait Video Segmentation
Weichen Xu, Yezhi Shen, Qian Lin, Jan P. Allebach, Fengqing Maggie Zhu · 2023
Portrait video segmentation has been widely used in applications such as online conferencing and content creation. However, it is challenging for mobile devices with limited computation resources to achieve accurate and temporal consistent real-time portrait segmentation. In this work, we propose a segmentation method based on the classic encoder-decoder architecture with a lightweight model design. To facilitate the efficient use of temporal guidance, our method takes an RGB-M input where M is a guidance portrait mask concatenated to the RGB input. Furthermore, we leverage the temporal guidance to enable model inference on the adaptive portrait region of interest (ROI). We introduce a two-stage training strategy to compensate for the limited data variety of portrait video datasets. Our method is evaluated on portrait videos including different types of daily activities, and outperforms existing portrait segmentation methods in terms of segmentation accuracy. Without introducing significant delay, our method is suitable for applications requiring real-time processing.