A study of automatic phonetic segmentation for forensic voice comparison
Chee Cheun Huang, Julien Epps · 2012
Forensic voice comparison (FVC) systems have often involved manual annotation of usable phonetic units, requiring substantial human labor. Recent research has shown the efficacy of automatic methods in FVC, and this paper investigates automatic phonetic segmentation in FVC systems. Nasals and vowels were found to contribute the most in terms of improvements in both the validity and reliability of the system. Results show that as a function of the duration of the recognized tokens there is a trade-off in which an improvement in validity corresponds to a degradation in reliability and vice versa. An implication is that minimizing the error of automatically estimated monophone boundaries may not necessarily result in the best system validity or reliability. A substantial improvement in log-likelihood-ratio cost (validity) of 17.02% and in 95% credible interval (reliability) of 5.97% over the baseline system was possible by fusing baseline scores with those from nasal and vowel segments.