Speech coding by limited weight neural networks (LWNN)
Bruno Gas, Jean‐Luc Zarader, P. Sellem, Jean-Charles Didiot · 2002
We present a new kind of speech coding. Usually the coding is obtained by a linear predictor LPC (or derivative LAR, LPCC) or by spectral analysis, as FFT, Cepstre or MECC. We propose to use a three layer neural network to learn phonemes extracted from the DARPA-TIMIT database. The network is designed to predict the next input signal value from the N previous ones. During the training stage, the first weight layer is the same for each phoneme. The second weight layer is different for each phoneme. In the generalization stage, the first weight layer remains fixed and initialized with those given by the training phase. When coding a test database phoneme, the output weight layer is trained to predict each phoneme values. The final neural predictive coding (NPC) corresponds to this second weight layer. We show that normalized coding can easily be obtained by using a nonlinear function of the weights instead of the weights themselves. Results are compared with others on temporal speech coding. A study of NPC by discriminant analysis and an application of MLP to phoneme recognition is presented.