Spanish Word Vectors from Wikipedia

Mathías Etcheverry, Dina Wonsever · 2016

Contents analisys from text data requires semantic representations that are difficult to obtain automatically, as they may require large handcrafted knowledge bases or manually annotated examples.Unsupervised autonomous methods for generating semantic representations are of greatest interest in face of huge volumes of text to be exploited in all kinds of applications.In this work we describe the generation and validation of semantic representations in the vector space paradigm for Spanish.The method used is GloVe (Pennington et al., 2014), one of the best performing reported methods , and vectors were trained over Spanish Wikipedia.The learned vectors evaluation is done in terms of word analogy and similarity tasks (Pennington et al., 2014;Baroni et al., 2014; Mikolov et al., 2013a).The vector set and a Spanish version for some widely used semantic relatedness tests are made publicly available.

Read the paper · More papers on PaperTik