EXPERIMENTS IN SYNTACTIC AND SEMANTIC CLASSIFICATION AND DISAMBIGUATION USING BOOTSTRAPPING
Robert P. Futrelle, Susan Gauch · 1993
Methods that generate word classes without requiring pretagging have had notable success in the last few years (bootstrap methods or unsupervised classification). The methods described here strengthen these approaches and produce excellent word classes from a 200,000 word corpus. The method uses mutual information measures plus positional information from the words in the immediate context of a target word to compute similarities. Using the similarities, classes are built using hierarchical agglomerative clustering. At the leaves of the classification tree, words are grouped by syntactic and semantic similarity. Further up the tree, the classes are primarily syntactic. Once the initial classes are found they can be used to improve the classification of single word instances --- to do classic word tagging. This is done by expanding each context word of a target instance into a tightly defined class of similar words, a simset. The use of simsets is shown to increase the tagging accuracy from 83% to 92% for the forms "cloned" and "deduced".