Using linguistic information to improve the performance of vector-based semantic analysis

Magnus Sahlgren, David Swanberg · DSpace repository (University of Tartu) · 2001

In this paper, we will show that the performance of vector-based semantic analysis can be improved by considering basic linguistic structures in the data-- e.g. morphology. For this purpose, we have used a new method for vector-based semantic analysis that computes semantic word vectors based on distributed representations by means of random labeling of words in narrow context windows. This form of representation is more natural than previously reported techniques, and, as we will show, equivalent or even superior in performance when subjected to a standardized synonym test.

Read the paper · More papers on PaperTik