Audio-visual bimodal speaker identification in a smart environment
Yanxiang Chen, Liu Ming · 2010
A bimodal person identification system is described by combining speech and 2D face images in a smart environment.The audio only system was based on a newly proposed model-segment-based Gaussian Mixture Model.The visual only system was a face recognition module based on K-nearest neighbors classifier.Finally the audio-visual system fused the individual modalities at the scoring level through score normalization,modality weighting and combination.Experimental results indicate the effectiveness of the speaker modeling methods and the fusion scheme.