99.8 percent accuracy achieved on Peterson and Barney (1952) acoustic measurements

Michael Stokes · The Journal of the Acoustical Society of America · 2014

In 2012, a paper was presented (Reetz, 2012) discussing the lack of working phonemic models, which was an acknowledgment to an earlier presentation (Ladefoged, 2004) discussing 50 + years of phonetics and phonology. These presentations highlighted the successes in phonological research over the last 60 and 50 years, respectively, but both concluded that there is still no recognized working model of phoneme identification. This presentation will discuss the Waveform Model of Vowel Perception (Stokes, 2009) achieving 99.8% accuracy on the Peterson and Barney (1952) dataset using 30 conditional statements across all ten vowels produced by the 33 males (509/510 for the vowels identified by humans at 100%). These results replicate and improve on the 99.2% achieved across the vowels produced by the males in the Hillenbrand (1995) dataset (Stokes, 2011). As a logical progression, ELBOW was developed in 2013 using the algorithm developed for static data to identify streaming vowel productions achieving over 91% before introducing improvements. Beyond ELBOW, it was essential to replicate earlier results on the most cited dataset in the literature. The Waveform Model has now replicated human performance across multiple datasets and is being successfully introduced into automatic speech recognition applications.

Read the paper · More papers on PaperTik