Robust Lip-Motion Features For Speaker Identification
H. Ertan Çetingül, Y. Yemez, Engin Erzin, Ahmet Murat Tekalp · 2006
The paper addresses the selection of robust lip-motion features for an audio-visual open-set speaker identification problem. We consider two alternatives for initial lip motion representation. In the first alternative, the feature vector is composed of the 2D-DCT coefficients of the motion vectors estimated within the detected rectangular mouth region, whereas in the second, lip boundaries are tracked over the video frames, and only the motion vectors around the lip contour are taken into account along with the shape of the lip boundary. Experimental results of the HMM-based identification system are included for performance comparison of the two lip motion representation alternatives.