Structural linear model-space transformations for speaker adaptation

Driss Matrouf, Olivier Bellot, Pascal Nocéra, Georges Linarès, Jean-François Bonastre · 2003

Within the framework of speaker-adaptation, a technique based on tree structure and the maximum a posteriori criterion was proposed (SMAP). In SMAP, the parameters estimation, at each node in the tree is based on the assumption that the mismatch between the training and adaptation data is a Gaussian PDF which parameters are estimated by using the Maximum Likelihood criterion. To avoid poor transformation parameters estimation accuracy due to an insufcienc y of adaptation data in a node, we propose a new technique based on the maximum a posteriori approach and PDF Gaussians Merging. The basic idea behind this new technique is to estimate an afne transformations which bring the training acoustic models as close as possible to the test acoustic models rather than transformation maximizing the likelihood of the adaptation data. In this manner, even with very small amount of adaptation data, the parameters transformations are accurately estimated for means and variances. This adaptation strategy has shown a signicant performance improvement in a large vocabulary speech recognition task, alone and combined with the MLLR adaptation.

Read the paper · More papers on PaperTik