MAP Based Speaker Adaptation in Very Large Vocabulary Speech Recognition of Czech
Petr Červa, Jan Nouza · Brno University of Technology Digital Library (Brno University of Technology) · 2004
Abstract. The paper deals with the problem of efficient adaptation of speech recognition systems to individual users. The goal is to achieve better performance in specific applications where one known speaker is expected. In our approach we adopt the MAP (Maximum A Posteriori) met-hod for this purpose. The MAP based formulae for the adaptation of the HMM (Hidden Markov Model) parame-ters are described. Several alternative versions of this met-hod have been implemented and experimentally verified in two areas, first in the isolated-word recognition (IWR) task and later also in the large vocabulary continuous speech recognition (LVCSR) system, both developed for the Czech language. The results show that the word error rate (WER) can be reduced by more than 20 % for a speaker who pro-vides tens of words (in case of IWR) or tens of sentences (in case of LVCSR) for the adaptation. Recently, we have used the described methods in the design of two practical appli-cations: voice dictation to a PC and automatic transcrip-tion of radio and TV news.