Consonant recognition methods for unspecified speakers using bpf powers and time sequence of LPC cepstrum coefficients
Katsuyuki Niyada, Masakatsu Hoshimi · Systems and Computers in Japan · 1987
Abstract This paper discusses the recognition of the consonant (except the consonant at the top of the word) in a word for unspecified speakers. First, the consonant section is detected based on the power dips extracted from the low‐ and high‐frequency power information, together with the nasal and unvoiced properties. Applying the discrimination diagram to the detected low‐ and high‐frequency dips, the phoneme is classified into the four phoneme classes (rough classification). Then methods are discussed which discriminate the individual phonemes in the phonemes group by pattern matching (fine classification). Using the time‐series pattern of the LPC cepstrum coefficient as the parameter, it is shown that the comparison with the standard phoneme patterns using the Bayes' discriminant and the Mahalabinos distance is the most useful. The result of recognition experiment using the segmentation, rough classification, and fine classifications is presented. For twenty subjects of both sexes, the mean recognition rate of 78.1 percent was achieved.