Extracting WordNet-like Top Concepts from Explanatory Dictionaries *

Hiram Calvo, Alexander F. Gelbukh · 2004

Abstract: Correct interpretation of the text frequently requires knowledge of semantic categories of nouns, especially in languages with free word order. For example, in Spanish the phrases pintó un cuadro un pintor (lit. painted a picture a painter) and pintó un pintor un cuadro (lit. painted a painter a picture) mean the same: ‘a painter painted a picture’; with the only way to tell the subject from the object being by knowing that pintor ‘painter ’ is causal agent cuadro is a thing. We present a method for extracting semantic information of this kind from existing machine-readable human-oriented explanatory dictionaries. First, we extract from the dictionary an is-a hierarchy and manually mark the categories of a few top-level concepts. Then, for a given word, we follow the hierarchy upward until finding a concept whose semantic category is known. Application of this procedure to two different human-oriented Spanish dictionaries gives additional information as compared with using solely Spanish EuroWordNet. In addition, we show the results of an experiment conducted to evaluate the similarity of word classification with this method. 1.

Read the paper · More papers on PaperTik