Simultaneous speaker normalisation and utterance labelling using Bayesian/neural net techniques
Simón Cox, John S. Bridle · International Conference on Acoustics, Speech, and Signal Processing · 2002
A particular form of neural network is described which has terminals for acoustic patterns, class labels, and speaker parameters. A method of training this network to tune in the speaker parameters to a new speaker is outlined. This process can also be viewed from a Bayesian perspective as maximizing the likelihood of the speaker's data by optimizing the model and speaker parameters. A method for doing this when the data are labeled is described. Results of using this technique with whole-word hidden Markov models (HMMs) indicate an improvement over speaker-independent performance and, for unlabeled data, a performance close to that achieved on labeled data.>