Discrimination of stop consonants using a data-driven analysis
Hitoshi Iwamida, Shinta Kimura · The Journal of the Acoustical Society of America · 1988
The characteristics of stop consonants are time-varying. In traditional running spectrum analysis, the frames are not always synchronized with the events of speech. In this presentation, a new data-driven analysis method is proposed in which the frames are synchronized with the events to extract the features of the stop consonants accurately. In this method, a feature vector is extracted from four or five frames that are equally spaced in a consonant segment. A voiced stop segment is defined as the segment between the release and the point where power exceeds a threshold. A voiceless stop segment is defined as the segment between the release and the voice onset. In these discrimination experiments, 94.0% of the voiced and 96.3% of the voiceless stops were correctly discriminated. The speech database used for these experiments was the Japanese monosyllables (/b,d,g,p,t,k/ + /a,i,u,e,o/) uttered by 20 speakers. It was confirmed that analysis synchronizing with consonant segments is effective for stop consonant discrimination.