Audio-visual person recognition: an evaluation of data fusion strategies

Claude C. Chibelushi · 1997

Audio-visual person recognition promises higher recognition accuracy than recognition in either domain in isolation. To reach this goal, special attention should be given to the strategies for combining the acoustic and visual sensory modalities. This paper presents a comparative assessment of three decision-level data fusion techniques for person identification. Under mismatched training and test noise-conditions, Bayesian inference and Dempster-Shafer theory are shown to outperform possibility theory. For these mismatched noise conditions, all three techniques result in compromising integration. Under matched training and test noise-conditions, the three techniques yield similar error rates approaching the more accurate of the two sensory modalities, and show signs of leading to enhancing integration at low acoustic noise levels. The paper also shows that automatic ident8cation of idenlicul twins is possible, and that lip margins convey a high level of speaker identity information

Read the paper · More papers on PaperTik