Forensic speaker recognition: Comparison and validation of three generations of automatic systems.
Didier Meuwly · The Journal of the Acoustical Society of America · 2010
The first aim of this paper is to demonstrate the improvement of automatic speaker recognition systems used for forensic evaluation over a period of 12 years. The second aim consists of exploring how the results of different systems can be compared and their improvement measured. The same set of experiments is replicated on three different systems: the original LR-based ASR system from 1998 (ASPIC I), the EPFL ASR system of 2004 (ASPIC II), and the Agnitio ASR system (BATVOX) from 2010. The reference database, Polyphone, consisting of 2000 male and 2000 female speakers from the French part of Switzerland, and the forensic database, Polyphone-IPSC, consisting of 16 male and 16 female speakers from the same region, were used to test the following forensic conditions: spontaneous speech, disguised speech, PSTN, GSM, signal to noise ratio from 30 to −3 dB, digital and analog recording, and close set of family related speakers. The results are visualized using Tippett plots. The refinement and calibration of the systems are measured and compared using the log-likelihood-ratio cost (Cllr). Finally, the value, in terms of forensic validation, of the methodology used and the results produced are discussed.