Automated induction of a lexical sublanguage grammar using a hybrid system of corpus- and knowledge-based techniques

G. Jan Wilms · 1996

Porting a Natural Language Processing (NLP) system to a new domain remains one of the bottlenecks in syntactic parsing, because of the amount of effort required to fix gaps in the lexicon, and to attune the existing grammar to the idiosyncracies of the new sublanguage. This dissertation shows how the process of fitting a lexicalized grammar to a domain can be automated to a great extent by using a hybrid system that combines traditional knowledge-based techniques with a corpus-based approach. An existing broad-coverage lexicon, the product of the expertise of seasoned linguists, is exploited to induce subcategorization features for new words, based on their paradigmatic relatedness to known words. The arbiter of paradigmatic relatedness is a category space, which can be bootstrapped from co-occurrence counts of content words in a training corpus. This dissertation uses a fixed-window approach which has been augmented with phrasal boundary information, and which has been finetuned by part-of-speech disambiguation of the input tokens. A smoothing technique called Singular Value Decomposition has been used to generalize the distributional information. Proximity in this reduced space is then used to find for all the context digests a neighborhood of words that are paradigmatically related; the subcategorization frames of a word are a composite of the features associated with these similar words. Experiments with PUNDIT, a broad-coverage symbolic NLP system, have shown that the category space can successfully be used to induce features like transitivity and subcategorization for clauses and infinitival complements. The data-driven process not only expands the lexicon for new words, it also fits the grammar to the new domain by adjusting the feature set of the existing verbs, adding object options where appropriate to increase coverage and removing them to purge unwanted false positives from the solution space. The advantage of combining data-driven mining with the existing lexical knowledgebase over other bootstrapping methods is that this approach does not require the manual identification of appropriate cues for subcategorization features, or the involved construction of a pattern matcher that is sophisticated enough to ignore false triggers.

Read the paper · More papers on PaperTik