Extending a thesaurus by classifying words

Takenobu Tokunaga, Atsushi Fujii, Sakurai Naoyuki, Hozumi Tanaka · 1997

This paper proposes a method for extending an existing thesaurus through classification of new words in terms of that thesaurus. New words are classified on the basis of relative probabilities of a word belonging to a given word class, with the probabilities calculated using nounverb co-occurrence pairs. Experiments using the Japanese Bunruigoihyo thesaurus on about 420,000 co-occurrences showed that new words can be classified correctly with a maximum accuracy of more than 80%. 1 Introduction For most natural language processing (NLP) systems, thesauri comprise indispensable linguistic knowledge. Roget's International Thesaurus [Chapman, 1984] and WordNet [Miller et al., 1993] are typical English thesauri which have been widely used in past NLP research [Resnik, 1992; Yarowsky, 1992]. They are handcrafted, machine-readable and have fairly broad coverage. However, since these thesauri were originally compiled for human use, they are not always suitable for computer-based na...

Read the paper · More papers on PaperTik