Consonant landmarks: Automatic detection and interpretation.

Chiyoun Park, Nancy F. Chen · The Journal of the Acoustical Society of America · 2008

Consonant landmarks are acoustic discontinuities in the speech signal that correspond to the closures and releases in speech production, and have been proposed as critical elements in speech processing [Stevens (2002)]. The three types of consonant landmarks represent the onset and offset of salient acoustic events: glottal vibration, turbulence noise, and sonorancy (e.g., nonvocalic voicing). While earlier work [Liu, (1996)] evaluated the success of identifying single candidate landmarks of all three types, this work focuses on two tools for evaluating strings of landmark candidates. First, a bigram model representing the physiologically feasible sequences of consonant landmarks is used to evaluate candidate strings. Second, a graphical method is used to identify the regions where the landmarks are reliably detected versus where they are ambiguous. Together these tools substantially improve the performance for landmark detection, identify regions in need of further acoustic analysis, and model the ‘grammatical’ structure of landmark sequences. Furthermore, the reliable regions in the proposed representation often correspond to structural elements such as lexical stress and word boundaries. Thus, the proposed representation is potentially useful in analyzing speech not only at the phoneme level but also at the word and phrase levels. [This work was supported by NIH/NIDCD DC02978 and T32DC00038.]

Read the paper · More papers on PaperTik