A comparison of feature representations for speaker-independent voiced-stop-consonant recognition
Bobby D. Bryant, J.N. Gowdy · 2002
The authors investigated various feature representations of speech which seem to provide robust estimates of machine-recognition-relevant parameters for the voiced-stop-consonant phoneme class. Instances of a block-windowed neural network (BWNN) were trained and tested using feature vectors extracted from data of up to four dialect regions in the TIMIT database. Three feature representations were chosen for use in this research based on their past performance in consulted feature representation studies. It is concluded that the feature representations produced by Seneff's (1988) auditory model particularly the mean-rate response representation, are good representations for voiced-stop consonant speech as well as vowel speech. It is also concluded that the addition of dynamic feature information in the form of differenced cepstral coefficients to the conglomerate mel-cepstral representative vectors made a difference in the recognition rate for voiced-stop consonants over the use of the mel-frequency cepstral coefficients alone. It can be hypothesized that the use of the BWNN architectures produced better recognition results than the use of other architectures that do not take into account the time and frequency variabilities encountered in utterances from different speakers.>