Detecting Polysemy in Hard and Soft Cluster Analyses of German Preposition Vector Spaces
Sylvia Springorum, Sabine Schulte im Walde, Jason Utt · 2013
This paper presents a methodology to identify polysemous German prepositions by exploring their vector spatial proper-ties. We apply two cluster evaluation metrics (the Silhouette Value (Kaufman and Rousseeuw, 1990) and a fuzzy ver-sion of the V-Measure (Rosenberg and Hirschberg, 2007)) as well as various cor-relations, to exploit hard vs. soft cluster analyses based on Self-Organising Maps. Our main hypothesis is that polysemous prepositions are outliers, and thus repre-sent either (i) singletons or (ii) marginals of the clusters within a cluster analysis. Our analyses demonstrate that (a) in a sub-set of the clusterings, singletons have a tendency to contain polysemous preposi-tions; and (b) misclassification and cluster membership rate exhibit a moderate corre-lation with ambiguity rate. 1