Generalizing a Strongly Lexicalized Parser using Unlabeled Data
Tejaswini Deoskar, Christos Christodoulopoulos, Alexandra Birch, Mark J. Steedman · 2014
Statistical parsers trained on labeled data suffer from sparsity, both grammatical and lexical.For parsers based on strongly lexicalized grammar formalisms (such as CCG, which has complex lexical categories but simple combinatory rules), the problem of sparsity can be isolated to the lexicon.In this paper, we show that semi-supervised Viterbi-EM can be used to extend the lexicon of a generative CCG parser.By learning complex lexical entries for low-frequency and unseen words from unlabeled data, we obtain improvements over our supervised model for both indomain (WSJ) and out-of-domain (questions and Wikipedia) data.Our learnt lexicons when used with a discriminative parser such as C&C also significantly improve its performance on unseen words.