A bootstrap training approach for language model classifiers
Volker Warnke, Elmar Nöth, Jan Buckow, Stefan Harbeck, Heinrich Niemann · 1998
In this paper, we present a bootstrap training approach for lan-guage model (LM) classifiers. Training class dependent LM and running them in parallel, LM can serve as classifiers with any kind of symbol sequence, e.g., word or phoneme sequences for tasks like topic spotting or language identification (LID). Irre-spective of the special symbol sequence used for a LM classifier, the training of a LM is done with a manually labeled training set for each class obtained from not necessarily cooperative speak-ers. Therefore, we have to face some erroneous labels and devia-tions from the originally intended class specification. Both facts can worsen classification. It might therefore be better not to use all utterances for training but to automatically select those utter-ances that improve recognition accuracy; this can be done by a bootstrap procedure. We present the results achieved with our best approach on the VERBMOBIL corpus for the tasks dialog act classification and LID. 1.