Combined optimisation of baseforms and model parameters in speech recognition based on acoustic subword units
Trym Holter, Torbjørn Karl Svendsen · 2002
A major challenge in speech recognition is creating a lexicon which is robust to inter and intra speaker variations. This is even more so in speech recognisers based on non linguistic units, e.g., acoustic subword units (ASWUs), since no standard pronunciation dictionaries are available. Thus the baseforms describing the vocabulary words in terms of the recognition units need to be generated from training data. We propose an algorithm for ASWU based speech recognition which performs a combined optimisation of the baseforms and the subword models. The resulting system has been tested on the DARPA Resource Management task, and is shown to perform comparably to a baseline phoneme based system.