Speaker Verification for Multi-Task Interactions
Y. Cai, X. Li, Z. Gong, T. R. Codina · Interacting with Computers · 2013
Human–computer interactions often involve multiple tasks such as navigation, searching, conversation and data retrieving. It is a challenge to identify the user in multi-tasking situations. Here, we present the algorithm for continuous verification of pre-defined users in both text-independent and text-dependent interactions. In the system, pitch features and Mel-Frequency Cepstral Coefficients are used for voice feature extraction. Vector quantization (VQ) and Mahalanobis distance are used for voice classification. We implemented our algorithms on the Android HTC phone. Based on our experimental results, we found that our training time can be within 1 min and the test time can be within 10 s or two sentences. VQ is suitable for a quick text-dependent verification such as voice signature; Mahalanobis distance, on the other hand, is better for continuous verification such as user authentication via voice. The empirical study also shows the impact of the verification time and languages on the verification accuracy. Our empirical study suggests that speaker verification algorithms work better when a user speaks in his or her native language.