Polarity Inducing Latent Semantic Analysis

Wen-tau Yih, Geoffrey Zweig, John C. Platt · 2012

Existing vector space models typically map synonyms and antonyms to similar word vec-tors, and thus fail to represent antonymy. We introduce a new vector space representation where antonyms lie on opposite sides of a sphere: in the word vector space, synonyms have cosine similarities close to one, while antonyms are close to minus one. We derive this representation with the aid of a thesaurus and latent semantic analysis (LSA). Each entry in the thesaurus – a word sense along with its synonyms and antonyms – is treated as a “document, ” and the resulting doc-ument collection is subjected to LSA. The key contribution of this work is to show how to as-sign signs to the entries in the co-occurrence matrix on which LSA operates, so as to induce a subspace with the desired property. We evaluate this procedure with the Grad-uate Record Examination questions of (Mo-hammed et al., 2008) and find that the method improves on the results of that study. Further improvements result from refining the sub-space representation with discriminative train-ing, and augmenting the training data with general newspaper text. Altogether, we im-prove on the best previous results by 11 points absolute in F measure. 1

Read the paper · More papers on PaperTik