EVALution 1.0: an Evolving Semantic Dataset for Training and Evaluation of Distributional Semantic Models

Enrico Santus, Frances Yung, Alessandro Lenci, Chu‐Ren Huang · 2015

In this paper, we introduce EVALution 1.0, a dataset designed for the training and the evaluation of Distributional Semantic Models (DSMs).This version consists of almost 7.5K tuples, instantiating several semantic relations between word pairs (including hypernymy, synonymy, antonymy, meronymy).The dataset is enriched with a large amount of additional information (i.e.relation domain, word frequency, word POS, word semantic field, etc.) that can be used for either filtering the pairs or performing an in-depth analysis of the results.The tuples were extracted from a combination of ConceptNet 5.0 and Word-Net 4.0, and subsequently filtered through automatic methods and crowdsourcing in order to ensure their quality.The dataset is freely downloadable 1 .An extension in RDF format, including also scripts for data processing, is under development.

Read the paper · More papers on PaperTik