Effective acoustic adaptation for a distant-talking interactive TV system
Jing Huang, Mark Epstein, Marco Matassoni · 2008
In this paper we have studied how to adapt a close-talking baseline acoustic model to a distant-talking application developed in an interactive TV dialogue system: distant-talking interfaces for control of interactive TV (DICIT) project. We have shown that in order to have effective adaptation from the outof-domain data it is better to acquire that data in the same DICIT environment than using contaminated data. By measuring grammar error rate (GER) and action classification error rate (AER) in addition to word error rate (WER), we have shown the best way to adapt the baseline model using available out-of-domain adaptation data (TIMIT) and small amount of in-domain (DICIT) adaptation data. The best approach is to use cascading MAP adaptation. With less than hours of out-ofdomain data and hour of in-domain data, the cascading MAP improves WER/GER/AER by / / relative respectively over the baseline model. The experimental results show that in-domain adaptation data is definitely needed to improve GER and AER. Index Terms: acoustic model adaptation, distant-talking speech recognition, dialog system