Bi-layer segmentation from stereo video sequences by fusing multiple cues
Yi Wu, Patricia P. Wang, Jianguo Li · 2008
Bi-layer video segmentation (segmentation of videos into foreground layer and background layer) has attracted a lot of research interests recently. The algorithm can be applied into many vision and multimedia applications, such as human computer interaction, gesture recognition, object detection/tracking and personal video editing. Traditional approaches cluster pixels into homogenous regions based on color distributions or motion patterns. However, spatial or temporal clues alone are insufficient to distinguish objects because different objects may share similar colors or motions. Nowadays, the availability of stereo cameras provides the possibility to recover certain 3D depth information from videos. In this paper, we propose bi-layer segmentation from stereo video sequences by fusing multiple cues, including 3D depth information, color/texture distribution, and motion vectors. We also explore temporal-spatial coherence among a consecutive sequence of frames in order to reduce noises from a single frame.