Reinforcing Parser Preferences through Tagging
Robbert Prins, van Gerardus Noord · 2003
Lexical ambiguity is an important source of inefficiency for wide-coverage HPSG parsing. In this paper, we propose a lexical analysis filter which removes unlikely lexical categories. The filter is implemented as a straightforward HMM n-gram POS-tagger, which computes the 'a posteriori' probability of each lexical category. A lexical category is removed if a competing lexical category is sufficiently more likely. The novel aspect of our approach is the fact that the tagger is trained on the output of the parser itself; therefore there is no need for hand-annotated material. Use of this filter increases the speed of the parser considerably, and in addition gives rise to an improvement in parsing accuracy.