Automating phonetic measurement: The case of voice onset time
Neville Ryant, Jiahong Yuan, Mark Yoffe Liberman · The Journal of the Acoustical Society of America · 2013
Of 58 papers published so far this year in Journal of Phonetics, 16 (28%) feature Voice Onset Time (VOT) or related measurements, confirming that VOT remains a central concern in the field. However, phoneticians’ VOT measurements generally continue to rely on human judgment, which requires significant labor, makes even large laboratory experiments onerous, and prevents the field from taking full advantage of the millions of hours of digital speech now becoming available. We present an algorithm for accurate automatic measurement of VOT, combining HMM forced alignment for determining approximate stop boundaries with paired burst and voicing onset detectors. Each detector is a frame-level max margin classifier operating on the scale-space projection of a small number of relevant acoustic features. On a large set of clean lab speech, this system has a mean absolute error (relative to human annotation) of only 2.8 ms, with 98% of errors <10 ms. On a subcorpus independently annotated by two of the authors, the system agreed with the two human annotators as well as they agreed with one another (1.49 vs 1.50 ms). Promising results on other datasets will be reported. The system will be released as open-source software.