A human benchmark for automatic speaker recognition

Marten van Dijk, Rosemary Orr, David van der Vloed, David A. van Leeuwen · 2013

Automatic Speaker Recognition has a potential to be used in Forensic Speaker Comparison. For the latter, forensic scientists agree that presentation of the comparison to court should be in terms of a calibrated likelihood ratio. In recent years the field of automatic speaker recognition has made significant progress in the analysis, evaluation and calibration of likelihood ratios. In this paper we investigate if speaker comparison by humans can be carried out using the same framework. For this, we use the US National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation 2010 material to measure the performance and calibrate the speaker comparison opinions of human subjects. Because empirical calibration needs a large collection of trials and a human judgment takes a substantial effort, the analysis is carried out for a collection of 40 subjects. From NIST SRE 2010 a subset of 1280 speaker comparison trials are selected. The selection is made using the scores from a state-of-the-art speaker recognition system, such that 1) the trials are representative of the overall performance, in terms of difficulty of the comparisons for the automatic system, and 2) they can be analyzed in three distinct classes ‘hard,’ ‘representative’ and ‘easy.’ Results show that this classification extends to the performance of the human collective, with an Equal Error Rate of 45 %, 25 % and 13 % respectively. Further, the overall human results can be calibrated using ROC convex hull analysis to show a nice linear relation between the 10-level similarity response and a log-likelihood-ratio scale.

Read the paper · More papers on PaperTik