Linguistic clues for corpus-based acquisition of lexical dependencies
Cécile Fabre, Didier Bourigault, Allées Antonio Machado · 2001
1 SYNTEX: a tool for the extraction of lexical dependencies Disambiguation of prepositional phrase attachment is a crucial issue for all NLP applications that need textual resources enriched with syntactic knowledge. We are confronted to this problem in the process of designing SYNTEX, a shallow parser specialised in the extraction of lexical dependencies (such as adjective/noun, or verb/noun associations) from French technical corpora. These word-to-word associations will be used as material for the construction of semantic classes on a distributional basis. In this context, the first step towards the automatic discovery of such dependencies is to determine to which word a preposition must be attached. For example, in the phrase disséquer le plateau rocheux en chevron, taken from a corpus in the domain of geomorphology1, the preposition en may potentially be attached to any of the three words disséquer, plateau, rocheux, as verbs, nouns or adjectives (and also adverbs) may govern a prepositional phrase. As lexico-syntactic information is part of what we want to extract from the text, we cannot rely on prior lexical resources: it is our belief, based on in-depth studies of corpora from technical domains, that words exhibit idiosyncratic uses from one domain to the other, not only at the semantic level, but also regarding their syntactic properties. As a consequence, our parser relies as much as