Integrated person identification using voice and facial features

Claude C. Chibelushi · 1997

Real-world automatic person recognition requires a consistently high recognition accuracy which is difficult to attain using a single recognition modality. This paper addresses the issue of person identification accuracy resulting from the combination of voice and outer lip-margin features. An assessment of feature fusion - based on audio-visual feature vector concatenation, principal component analysis, and linear discriminant analysis - is conducted. The paper shows that outer lip margins carry speaker identity cues. It is also shown that the joint use of voice and lip-margin features is equivalent to an effective increase in signal-to-noise ratio of the audio signal. Simple audio-visual feature vector concatenation is shown to be an effective method for feature combination, and linear discriminant analysis is shown to possess the capability of packing discriminating audio-visual information into fewer coeficients than principal component analysis.

Read the paper · More papers on PaperTik