Algorithms for video object detection and segmentation with application to content-based multimedia systems
Alexandros Eleftheriadis, Huitao Luo · 2000
We have designed several analysis algorithms for video object (VO) detection and segmentation, which is an important part of the new MPEG4 and MPEG7 standards. The algorithms are developed in two directions, i.e., automatic algorithms and semi-automatic algorithms. In automatic analysis research, our work is mainly focused on model-based algorithm design. Instead of working on general-purpose algorithms, we design detection and segmentation algorithms for typical multimedia applications. By constraining the application domain, we could represent the video objects to be processed with some models. The models are created with specific statistical structures, but are still generally applicable in certain application domains. Specifically, we have designed video object detection and/or segmentation algorithms for three applications: (1) Real-time VO segmentation for videophones. (2) Anchorperson detection and segmentation for broadcast news indexing and retrieval. (3) Face detection in the compressed DCT domain. For videophone applications, we design a blob-based region model and a shape model to represent typical head-and-shoulder foreground. Both models use a Gaussian assumption for their feature vectors. At the system level, a hierarchical structure is designed to support an online processing of model-creating, model-fitting, and model-updating. In our experiments, a QCIF size video sequence is segmented in real time using software only into three video objects: a background, a head and a shoulder on average Pentium PC platforms. Based on the real time performance of the algorithm, we discuss two direct applications of it, i.e., real time VO generation for MPEGA codecs and content-based bit-rate control for traditional H.263 codecs. For anchorperson detection, we propose to model anchorperson patterns with their color and shape features. The detection problem is decomposed into a color model based face region detection and a shape model based head-and-shoulder pattern detection problem. The statistical shape model design is similar to what we used for videoconference applications, except an offline model fitting. Our work on face detection proposes to merge two types of face detection algorithms in the literature, skin-color based and texture based face detection. Though each of them has been explored extensively, our work for the first time shows that they can be combined with a hybrid statistical model (color-texture model) to generate better detection performance. In addition, the hybrid model is designed in the compressed DCT domain. A number of fundamental problems, e.g., block quantization problem, preprocessing, and feature vector selection and classification in the DCT domain, are discussed. In semi-automatic analysis research, an interactive authoring system is designed for video object segmentation. This system features a new contour interpolation algorithm, which enables the user to define the contour of a VO on multiple anchor frames while the computer interpolates the missing contours of this object on every frame automatically. Typical active contour model is adapted and the contour interpolation problem is decomposed into two directional contour tracking problems and a merging problem. In addition, new user interaction models are created for the user to interact with the computer. Experiments indicate that this system offers a good balance between algorithm complexity and user interaction efficiency.