Adaptive training for robust ASR

Mark Gales · 2005

Adaptive training is a powerful training technique for building speech recognition systems on nonhomogeneous data. The aim is to remove unwanted variability, such as changes in speaker, channel or acoustic environment, from desired changes, the acoustic differences between words. During training, two sets of models are generated: a canonical model set for the desired "true" variability of the speech data, and a set of transforms to represent the unwanted variability. The canonical model set trained in this fashion should be more "amenable" to being adapted to a particular target condition and more "compact". During recognition, a transform to the target domain is trained. This target specific transform is then used with the canonical model set in the recognition process. The paper gives an overview of the underlying theory and assumptions used in adaptive training. Furthermore, the use of adaptive training schemes in current state-of-the-art tasks is described, together with a discussion of how such schemes may be used in the future.

Read the paper · More papers on PaperTik