Preprocessing and neural classification of English stop consonants [b, d, g, p, t, k]

Anna Aurelia Esposito, C. E. Ezin, Michele Ceccarelli · 2002

Neural networks are accepted as powerful learning tools in pattern recognition in which they proved their performance. Nevertheless, many problems like phoneme classification with a multi-speaker continuous speech database are hard even for neural networks. The authors' aim is to propose an artificial neural network architecture that detects acoustic features in speech signals and classifies them correctly. They reached this goal with English stop consonants [b, d, g, p, t, k] extracted from the general multi-speaker database (TIMlT) by modifying some parameter values in the preprocessing algorithm and by using a modified TDNN (time delay neural network) architecture. The net performed a good classification giving as testing recognition percentage the following results: 92.9 for [b], 91.8 for [d], 92.4 for [g], 80.3 for [p], 90.2 for [t], 91.2 for [k].

Read the paper · More papers on PaperTik