Unsupervised speaker adaptation of spectra based on a minimum fuzzy vector quantization error criterion
Hiroshi Matsumoto, Yasuki Yamashita · The Journal of the Acoustical Society of America · 1988
Reference spectral code vectors are adapted to an input speaker by adding a distance-dependent linear combination (an interpolated vector) of speaker difference vectors (unknown adaptation vectors) at given typical points in the reference spectral space. The unknown adaptation vectors are estimated so as to minimize the total fuzzy vector quantization errors for training vectors using an adapted reference codebook. This minimization is iteratively resolved under a restriction on the sum of the norms over all the unknown adaptation vectors. The above adaptation process is applied in a stepwise manner by increasing the number of typical spectral points. The fuzziness, interpolation parameters and the limits of the norm were examined through 28 word recognition tests for 4 male and 4 female speakers, using reference patterns from a male speaker. Under the best conditions, with 16 adaptation vectors, this method achieved the same recognition accuracy as a supervised speaker adaptation for male speakers with training samples as short as 3 s.