An extended formulation of score-based fusion scheme for robust audio-visual speaker identification
Md. Tariquzzaman, Jin Young Kim, Seung You Na, Min Gyu Song · 한국정보기술학회논문지 · 2011
The score based integration scheme based on confidence measure(CM) derived from scores is widely used in audio-visual speaker/speech recognition. However, the conventional CM has the problems of missing-reliability and miss-estimation (overestimation or underestimation). In a typical audio-visual integration scheme each classifier generate a class of log-likelihood score which are first normalized to measure the confidence. The conventional confidence measuring approach uses the highest and second highest score values of the normalized scores. In this paper first, we have proposed a function through introducing thresholding parameter to account reliable second highest values of the scores thus to account reliable confidence. Secondly, a weighting parameter is introduced on the second highest values to control overestimation and underestimation problem in reliability measure and thus to account the optimal confidence for individual classifier simultaneously. To evaluate the proposed approach, we perform the speaker identification experiments using VidTimit database. In the experiments, acoustic noises from NOISEX92 database and artificial illumination are applied to the original test audio and visual signals respectively. The Experimental results show that our proposed approach have effectively enhanced the identification accuracy of about 1%~4% through correcting the problems of missing-reliability and miss-estimation in comparison with the baseline system at different noisy conditions in comparison with the baseline system.