Syntactic Annotation for the Spoken Dutch Corpus Project (CGN)
Heleen Hoekstra, Michael Moortgat, Ineke Schuurman, Ton van der Wouden · 2001
We present a computationally inexpensive technique to perform style adaptation: a general training text corpus is statically adapted to the style of a given target recognition task by weighted counting. The n-gram language model derived from this weighted background corpus is then used as a component of a mixture language model in a word lattice rescoring framework. We specify two types of weighted counting and evaluate their effectiveness in terms of word recognition error rate. Adapting broadcast news style to news talkshow style, these methods yield a close to insignificant improvement with respect to the unweighted case. However, adapting financial newspaper style to news talkshow style, we observed a 0.7% absolute reduction of word recognition error rate.