Exploiting the WWW as a corpus to resolve PP attachment ambiguities
Martin Volk · Zurich Open Repository and Archive (University of Zurich) · 2001
Introduction Finding the correct attachment site for prepositional phrases (PPs) is one of the hardest problems when parsing natural languages. An English sentence consisting of a subject, a verb, and a nominal object followed by a prepositional phrase is a priori ambiguous. The PP in sentence 1 is a noun attribute and needs to be attached to the noun, but the PP in 2 is an adverbial and thus part of the verb phrase. (1) Peter reads a book about computers. (2) Peter reads a book in the subway. If the subcategorisation requirements of the verb or the competing noun are known the ambiguity can sometimes be resolved. But many times there are no clear requirements. Therefore, there has been a growing interest in using statistical methods that reflect attachment tendencies. This new line of research was kicked off by Hindle and Rooth (1993). They tackled the PPattachment ambiguity problem (for English) by computing lexical association sco