Audio-visual bimodal speaker identification in a smart environment

Yanxiang Chen, Liu Ming · 2010

A bimodal person identification system is described by combining speech and 2D face images in a smart environment.The audio only system was based on a newly proposed model-segment-based Gaussian Mixture Model.The visual only system was a face recognition module based on K-nearest neighbors classifier.Finally the audio-visual system fused the individual modalities at the scoring level through score normalization,modality weighting and combination.Experimental results indicate the effectiveness of the speaker modeling methods and the fusion scheme.

Read the paper · More papers on PaperTik