Joint optimization of classifier and feature space in speech recognition
Gary M. Kuhn · 2003
The author presents a feedforward network which classifies the spoken letter names 'b', 'd', 'e', and 'v' with 88.5% accuracy. For many poorly discriminated training examples, the outputs of this network are unstable or sensitive to perturbations of the values of the input features. This residual sensitivity is exploited by inserting into the network a new first hidden layer with localized receptive fields. The new layer gives the network a few additional degrees of freedom with which to optimize the input feature space for the desired classification. The benefit of further, joint optimization of the classifier and the input features was suggested in an experiment in which recognition accuracy was raised to 89.6%.>