A Computational Theory of Lexical Relatedness
Marc Light · 1993
Lexicon coverage is often the limiting factor in natural language processing systems. Recent work has attempted to remedy this situation by extracting information from machine readable dictionaries. Unfortunately, no NLP lexicon system or dictionary could possibly list all the potential words of English. However, humans are often able to interpret novel word forms (that is, words they have not seen before) without difficulty. One way we do this, if the word is complex (e.g., "undecidability"), is by using cues from the internal structure of the word. Relations in phonological form often correspond to relations in meaning. For example, if someone knows what the verb "open" means, a number of educated guesses can be made about the meaning of "reopen". Exceptions abound in lexical data and any system that attempts to use lexical generalizations must be able to handle exceptions in a principled fashion. In this report, I will describe the preliminary design of a system that uses relations in form to derive relations in meaning. For a new word, the system will produce meaning postulates that represent an educated guess about the meaning of the new word. These meaning postulates will be written in Episodic Logic, and the entire system will be a module of the TRAINS system.