Speaker-independent features extracted by a neural network

Yohei Kato, Masashi Sugiyama · IEEE International Conference on Acoustics Speech and Signal Processing · 1993

The authors propose an algorithm using a neural network to normalize features that differ between speakers in speaker-independent speech recognition. The algorithm has three procedures: (1) initially training a neural network, (2) calculating the alignment function between the target signal and the network's output by dynamic time warping, and (3) incrementally training the network for extracting speaker-independent features. The neural network is a fuzzy partition model (FPM) with multiple input-output units to give a probabilistic formulation. The algorithm was evaluated in phrase recognition experiments by FPM-LR recognizers. The FPM was directly combined with a LR parser. The algorithm is compared with a conventional training algorithm in terms of recognition performance. The experimental results show that a neural network can be used as a new speaker-independent feature extractor.>

Read the paper · More papers on PaperTik