First-order and second-order context representations: geometrical considerations and performance in word-sense disambiguation and discrimination

Alfredo Maldonado, Martin Emms · 2012

First-order and second-order context vectors (C 1 and C 2 ) are two rival context representations used in word-sense disambiguation and other endeavours related to distributional semantics. C 1 vectors record directly observable features of a context, whilst C 2 vectors aggregate vectors themselves associated to the directly observable features of the context. Whilst C 2 vectors may appeal on a number of grounds, such as being less sparse and leveraging additional information from a larger corpus, not much work has been devoted to contrasting C 2 with C 1 vectors. While the concerns of the paper are primarily empirical we also advocate a particular formulation of C 2 vectors, whereby C 2 vectors (of dimensionality f2) are derived from C 1 vectors (of dimensionality f1) by post-multiplication by some f1 f2 matrix. This makes plainer the relation of the C 2 construction to standard methods for dimensionality reduction. We then consider two geometric properties of C 1 - and C 2 -based sense vectors for sense-tagged data. We show that, perhaps surprisingly, the C 2 -based representation of a sense is not to any great extent parallel (similar) to the C 1 -based representation of that sense. We also show that the angular spread amongst the C 1 -based sense vectors is considerably greater than the spread amongst the C 1 versions. Following on from this, we then compare both sense vectors in supervised word sense disambiguation and unsupervised word sense discrimination settings, finding the C 1 -based vectors superior to the C 2 -based vectors in the supervised setting, but quite similar in performance in the unsupervised setting.

Read the paper · More papers on PaperTik