Detection of consonant voicing: A module for a hierarchical speech recognition system
Jeung-Yoon Choi · The Journal of the Acoustical Society of America · 1999
This research describes a module for detecting consonant voicing in a hierarchical speech recognition system. In this system, acoustic cues are used to infer values of features that describe phonetic segments. A first step in the process is examining consonant production and conditions for phonation, to find acoustic properties that may be used to infer consonant voicing. These are examined in different environments to determine a set of reliable acoustic cues. These acoustic cues include fundamental frequency, difference in amplitudes of the first two harmonics, cutoff first formant frequency, and residual amplitude of the first harmonic, around consonant landmarks. Classification experiments are conducted on hand and automatic measurements of these acoustic cues for isolated and continuous speech utterances. Voicing decisions are obtained for each consonant landmark, and are compared with lexical and perceived voicing for the consonant. Performance is found to improve when measurements at the closure and release are combined. Training on isolated utterances gives classification results for continuous speech that is comparable to training on continuous speech. The results in this study suggest that acoustic cues selected by considering the representation and production of speech may provide reliable criteria for determining consonant voicing.