Vector-based semantic analysis: representing word meanings based on random labels

Magnus Sahlgren · 2001

Vector-based semantic analysis is the practice of using co-occurrence statistics to construct vectors that represent word meanings by virtue of their direction in multi-dimensional semantic space. This paper discusses the theoretical presumptions behind this practice, and a representational scheme based on the Distributional Hypothesis is identified as the rationale for vector-based semantic analysis. A new method for calculating semantic word vectors is then described. The method uses random labeling of words in narrow context windows to calculate semantic context vectors for each word type in the text data. The method is evaluated with a standardized synonym test, and it is shown that incorporating linguistic information in the context vectors can enhance the results.

Read the paper · More papers on PaperTik