Vector-based semantic analysis: representing word meanings based on random labels
Magnus Sahlgren · 2001
Vector-based semantic analysis is the practice of using co-occurrence statistics to construct vectors that represent word meanings by virtue of their direction in multi-dimensional semantic space. This paper discusses the theoretical presumptions behind this practice, and a representational scheme based on the Distributional Hypothesis is identified as the rationale for vector-based semantic analysis. A new method for calculating semantic word vectors is then described. The method uses random labeling of words in narrow context windows to calculate semantic context vectors for each word type in the text data. The method is evaluated with a standardized synonym test, and it is shown that incorporating linguistic information in the context vectors can enhance the results.