Comparison of human listeners and speaker verification systems using voice mimicry data

Ville Hautamäki, Rosa González Hautamäki, Tomi Kinnunen, Anne-Maria Laukkanen · 2014

In this work, we compare the performance of human listeners and two well known speaker verification systems in presence of voice mimicry. Our focus is to gain insights on how well human listeners recognize speakers when mimicry data is included and compare it to the overall performance of state-ofthe-art speaker verification systems, a traditional Gaussian mixture model-universal background model (GMM-UBM) and an i-vector based classifier withcosine scoring. Wehave found that for the studied material in Finnish language, the mimicry attack was able to slightly increase the error rate in a range acceptable for the general performance of the system (EERfrom 9 to 11%). Our data reveals that enhancing the audio material by minimizing the differences of data collected in different environments improves the accuracy of the speaker verification systems even in the presence of mimicked speech. The performance of the human listening panel shows that successfully imitated speech is difficult to recognize, even more difficult to recognize a person who is intentionally trying to modify his or her own voice. The average listener made 8 errors from 34 selected trials while the automatic systems had 6 error in the same set.

Read the paper · More papers on PaperTik