A fully-automated approach to creating computational semantic lexicons and its implementation framework

James Robert Slagle, Sait Dogru · 1996

Research on modern theories of grammar has focused on the lexical semantics as the most important aspect of a successful NLP system. One of the latest and most successful modern linguistic frameworks, Head-Driven Phrase Structure Grammar (HPSG), for instance, has only a handful of grammar rules, but depends almost entirely on the structure of its lexical entries through signs. A model of the human lexicon is now widely recognized as the single most important factor in the success of modern NLP systems. This lexical approach to NLP has produced the new generation of successful NLP applications. Consequently, the acquisition and representation of lexical information have become fundamental issues in modern computational linguistics. Unfortunately, most of these approaches are manual and as such, are burdened by several problems including acquisition inconsistencies. Statistical and semi-automated approaches, on the other hand, deal with the surface structure of the lexicon and ignore the semantics that individual lexical words represent. In this thesis we propose and implement a two-level theory of the human lexicon. Our theory addresses both the acquisition of the lexicon as well as the representation of the semantic content of lexical entries. We claim that the human lexicon has been traditionally viewed as a simple lookup table, ignoring the underlying semantic layer. We show how both of the levels can be naturally captured within the framework of our theory by analyzing definitions of words of all major categories of the English grammar, as found, for instance, in hard-copy or on-line dictionaries. This process starts by accepting sequences of words in paragraphs and tagging these words and identifying the major syntactic constituents. Each sentence in the paragraph is then analyzed, both semantically and within the contextual constraints of the previously analyzed sentences, to produce an internal semantic representation in terms of Conceptual Graphs, which are evaluated via our implementation of the formula translator. These graphs capture the denotational semantics of individual lexical words. We show how this two-level view of the lexicon enables one of the widest-coverage grammars of English and how it can be applied to novel problem domains.

Read the paper · More papers on PaperTik