The extraction method of the word meaning class

K. Tsuda, M. Nakamura · 2003

In natural language processing, the semantic class information about a word is an important piece of knowledge. For a thesaurus dictionary, which shows the semantic information between words, the editing work is normally carried out manually, which means that a great number of man-hours is necessary for the editing work. This paper proposes a method of extracting the semantic class information on a word from a set of documents. This information is extracted by using the characteristic that the frequency of abstract words is high while the frequency of concrete words is small. As a result of this experiment, it was confirmed that about 20% of the extracted words should be registered in the thesaurus dictionary.

Read the paper · More papers on PaperTik