Invariant acoustic cues of consonants in a vowel context

Riya Singh, Jont B. Allen · The Journal of the Acoustical Society of America · 2011

The classic JASA papers by French and Steinberg (1947), Fletcher and Galt (1950), Miller and Nicely (1955), and Furui (1986) provided us with detailed CV+VC confusions due to masking noise and bandwidth and temporal truncations. FS47 and FG50 led to the succinctly summarizing articulation index (AI), while MN55 first introduced information-theory. Allen and his students have repeated these classic experiments and analyzed the error patterns for large numbers of individual utterances [http://hear.beckman.illinois.edu/wiki/Main/Publications], and showed that the averaging of scores removes critical details. Without such averaging, consonant scores are binary, suggesting invariant features used by the auditory system to decode consonants in isolated CV. Masking a binary feature causes the consonant error to jump from zero to chance (within some small subgroup of sounds), with an entropy determined by conflicting cues, typically present in naturally spoken sounds. These same invariant features are also used when decoding sentences having varying degrees of context. A precise knowledge of acoustic features has allowed us to reverse engineer Fletcher's error-product rule (FG50), providing deep insight into the workings of the AI. Applications of this knowledge is being applied to a better understanding of the huge individual differences in hearing impaired ears and machine recognition of consonants.

Read the paper · More papers on PaperTik