Speaker adaptation based on two-step active learning

Koichi Shinoda, Hiroko Murakami, Sadaoki Furui · 2009

We propose a two-step active learning method for supervised speaker adaptation. In the first step, the initial adaptation data is collected to obtain a phone error distribution. In the second step, those sentences whose phone distributions are close to the error distribution are selected, and their utterances are collected as the additional adaptation data. We evaluated the method us- ing a Japanese speech database and maximum likelihood linear regression (MLLR) as the speaker adaptation algorithm. We confirmed that our method had a significant improvement over a method using randomly chosen sentences for adaptation. Index Terms: speech recognition ,speaker adaptation, active learning nition accuracies in comparison with the conventional adapta- tion frameworks when the initial speaker-independent acoustic model has already been tuned to the given recognition task. In this paper, we propose a two-step active learning method for supervised speaker adaptation. In the first step, our method collects a small amount of utterances from a user to obtain his/her tendency in speech recognition errors. In the second step, it selects those sentences rich in phonetic units in the er- rors from a sentence pool, and it collects their utterances as ad- ditional data for adaptation. Since our method directly aims at decreasing recognition errors, it is expected to be highly dis- criminative. Our method has two critical issues: One is how to relate the recognition errors to the selection criterion of adap- tation sentences. The other is how to set the size of the initial adaptation data in the first step; it should be as small as possi- ble in order to decrease the user's effort, but it must be sufficient enough to estimatethe tendency of errors precisely. Wedescribe our approach to resolving these issues in this paper. This paper is organized as follows. Section 2 explains our method, and Section 3 briefly explains the MLLR speaker adap- tation method. Section 4 reports our evaluation experiments us- ing a Japanese speech database, and Section 5 concludes the paper.

Read the paper · More papers on PaperTik