Discriminating English Word Senses Using Cluster Analysis
Paul Watters · Journal of Quantitative Linguistics · 2002
Recently, statistical models for the identification of word senses in English text have been suggested, such as Latent Semantic Analysis (LSA), which is based on dimensionality reduction. While this approach has yielded promising results, it makes many assumptions about the underlying semantic structure. In this paper, the goal is to use cluster analysis to group word senses objectively on the basis of their co-occurrence with other words. This method does not make any a priori assumptions about the group to which a case might be assigned: It is an arbitrary classification made on the basis of a specific number of a group, which are then classified on the basis of their metric distance from one another in a high-dimensional space. The results of classifying two senses of the word BANK indicate high classification accuracy for primary word senses, but poor classification accuracy for secondary word senses. A role for using cluster analysis to determine highly discriminating items in text is discussed.