Hierarchical probabilistic models for video object segmentation and tracking
David Thirde, Graeme Jones · 2004
The Problem: The goal of segmentation and tracking video objects in generic scenes is to segment the objects accurately and consistently depending on a set of semantics defined. Methods previously applied to this problem can be divided primarily into region-based or boundary-based methods. These two distinct approaches attempt to locate an object based on the semantic homogeneity of feature vector regions or by measuring gradient information in the feature space to locate object boundaries. Many techniques for object extraction allow a human operator to locate and define the semantic video objects to be segmented and tracked. Turning this problem into one of classification, the image data can be taken to be an array of feature vectors and a classifier can be used to assign the set of user provided labels to unlabelled feature vectors. The information contained in the feature vector and choice of classifier is dependent on the application and/or semantics defined. Motivation: Semantic video objects with regard to the human visual system often have a complex underlying probability density function within the image feature space and hence previous approaches to this problem have applied many existing parametric and non-parametric forms of representation to extract such objects on an accurate pixel-wise basis. Parametric models have previously been applied to locate and segment video objects on a per-frame basis (e.g. Noel and O'Connor), although the functional form of the density model may not always provide a good representation of the object PDF within the feature space and the resulting segmentation mask quality is often degraded as a result.