Speaker adaptation using maximum a posteriori probability estimation with data size-dependent parameter smoothing
Masahiro Tonomura, Tetsuo Kosaka, Shoichi Matsunaga · Systems and Computers in Japan · 1999
In this paper, a speaker adaptation method is proposed that can be used efficiently in a wide range of available adaptation data in the case of unspecified speaker model based on continuous density HMM. The proposed method combines maximum a posteriori probability estimation (MAP estimation, a method to estimate parameters using preformed knowledge) and transfer vector field smoothing (a method to interpolate and smooth parameters). This makes it possible to improve adaptation performance in the case of scarce data available for adaptation; at the same time, control of smoothing intensity in accord with the data size is implemented, which provides a single adaptation framework for a wide range of data size. Experiments, with only Japanese phoneme connectivity constraints imposed, proved that the proposed method outperforms MAP estimation with up to 20 phrases while achieving nearly the same performance with more data available for adaptation. © 1999 Scripta Technica, Syst Comp Jpn, 30(11): 59–66, 1999