The modified kanerva model: theory and results for real-time word recognition
Thomas J. Clarke, Richard W. Prager, F. Fallside · IEE Proceedings F Radar and Signal Processing · 1991
The use of the modified Kanerva model to perform word recognition is continuous speech after being trained on the multi-speaker Alvey ‘Hotel’ speech corpus is described. A theoretical analysis has enabled us to increase the speed of execution of part of the model by two orders of magnitude over that previously reported by Prager and Fallside. The memory required for the operation of the model has been similarly reduced. The recognition accuracy reaches 95% without syntactic constraints when tested on different data from the seven trained speakers. Real-time simulation of a model with 9734 active units is now possible in both training and recognition modes using the Alvey Parsifal transputer array. The modified Kanerva model is a static network consisting of a fixed nonlinear mapping (location matching) followed by a single layer of conventional adaptive links. A section of preprocessed speech is transformed by the nonlinear mapping to a high dimensional representation. From this intermediate representation a simple linear mapping is able to perform complex pattern discrimination to form the output, indicating the nature of the speech features present in the input window. The major advantage of this architecture lies in the speed and robustness of iterative algorithms available for single layer networks. They are faster and more stable than schemes involving error back propagation and, with a suitable nonlinear preprocessor, can compute complex discriminant functions.