Which Statistics Reflect Semantics? Rethinking Synonymy and Word Similarity
Derrick Higgins · 2005
viewed as an engineering task (with the notable exception of much writing on Latent Semantic Analysis (LSA)), the relative success of different approaches to constructing word similarity measures is highly relevant to issues in theoretical semantics and language acquisition. With this background in mind, this paper has two main aims. First, we will present yet another statistical approach to the calculation of word-similarity scores (LC-IR), which significantly outperforms other methods on standard benchmarks including the 80-question set of TOEFL ® synonym test items first employed by Landauer and Dumais (1997). 1 Second, we hope to demonstrate that • various methods for assessing word similarity are based on fundamentally different assumptions about the statistical properties which synonyms can be expected to display, • the performance of each method can be taken as a judgment on the validity of these assumptions, and • whether these predictions regarding the statistical distribution of synonyms in a corpus are borne out ought to be taken into account in any consideration of the acquisition of meaning as part of language, and the mental representation of meaning. 1 2 Derrick Higgins 2 Statistical approaches to word similarity Without indulging in too much of a caricature, we can classify different approaches to statistical estimation of word similarity according to the assumptions which they make about the distribution of synonyms (actually, plesionyms; cf. Edmonds and Hirst (2002)). The three main assumptions made by existing word similarity measures are the topicality assumption, the proximity assumption, and the parallelism assumption. 2.1 Topicality: LSA et al. The techniques of Latent Semantic Analysis, Random Indexing, and Lund