Non-Linear Predictors based on the Functionally Expanded Neural Networks for Speech Feature Extraction
Mohamed Chétouani, Amir Hussain, Bruno Gas, Jean‐Luc Zarader · 2006
In this paper we focus on the design of the feature extractor stage of the speech recognition system which aims to compute optimal vectors for the next phoneme classification stage. We propose a new non-linear feature extraction method based on the linear-in-parameters Functionally Expanded Neural Network (FENN) model. The main idea is to design an improved and flexible feature extractor which can effectively account for some of the significant non-linear phenomena usually observed in the speech production process. The effectiveness of the proposed method is assessed on phoneme classification tasks. Specifically, we evaluate the performances on the telephone quality NTIMIT database, focusing the investigations on highly confusable phonemes such as front vowels: /ih/, /ey/, /eh/, /ae/. The results are compared with other widely used coding methods namely, the Linear Predictive Coding (LPC) and the Mel Frequency Cepstral Coding (MFCC). The experiments show a relative improvement in the rates through the use of our proposed non-linear feature extractor technique.