Clustering contextual facial display sequences

Jesse Hoey · 2003

Describes a method for learning classes of facial motion patterns from a video of a human interacting with a computerized embodied agent. The method also learns correlations between the discovered motion classes and the current interaction context. Our work is motivated by two hypotheses. First, a computer user's facial displays are context-dependent, especially in the presence of an embodied agent. Second, each interactant uses their face in different ways, for different purposes. Our method describes facial motion using optical flow over the entire face, projected to the complete orthogonal basis of Zernike polynomials. A context-dependent mixture of hidden Markov models (cmHMM) clusters the resulting temporal sequences of feature vectors into facial display classes. We apply the clustering technique to sequences of continuous video, in which a single face is tracked and spatially segmented. We discuss the classes of patterns discovered for a number of subjects.

Read the paper · More papers on PaperTik