Domain Adaptation with Unlabeled Data for Dialog Act Tagging
Anna Margolis, Karen Livescu, Mari Ostendorf · 2010
We investigate the classification of utterances into high-level dialog act categories using word-based features, under conditions where the train and test data differ by genre and/or language. We handle the cross-language cases with machine translation of the test utterances. We analyze and compare two featurebased approaches to using unlabeled data in adaptation: restriction to a shared feature set, and an implementation of Blitzer et al.’s Structural Correspondence Learning. Both methods lead to increased detection of backchannels in the cross-language cases by utilizing correlations between backchannel words and utterance length. 1