SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection
Dominik Schlechtweg, Barbara C. McGillivray, Simon Hengchen, Haim Dubossarsky, Nina Tahmasebi · 2020
Lexical Semantic Change detection, i.e., the task of identifying words that change meaning over time, is a very active research area, with applications in NLP, lexicography, and linguistics.Evaluation is currently the most pressing problem in Lexical Semantic Change detection, as no gold standards are available to the community, which hinders progress.We present the results of the first shared task that addresses this gap by providing researchers with an evaluation framework and manually annotated, high-quality datasets for English, German, Latin, and Swedish.33 teams submitted 186 systems, which were evaluated on two subtasks. OverviewRecent years have seen an exponentially rising interest in computational Lexical Semantic Change (LSC) detection (Tahmasebi et al., 2018;Kutuzov et al., 2018).However, the field is lacking standard evaluation tasks and data.Almost all papers differ in how the evaluation is performed and what factors are considered in the evaluation.Very few are evaluated on a manually annotated diachronic corpus (McGillivray et al., 2019;Perrone et al., 2019; Schlechtweg et al., 2019, e.g.).This puts a damper on the development of computational models for LSC, and is a barrier for high-quality, comparable results that can be used in follow-up tasks.We report the results of the first SemEval shared task on Unsupervised LSC detection.1 We introduce two related subtasks for computational LSC detection, which aim to identify the change in meaning of words over time using corpus data.We provide a high-quality multilingual (English, German, Latin, Swedish) LSC gold standard relying on approximately 100,000 instances of human judgment.For the first time, it is possible to compare the variety of proposed models on relatively solid grounds and across languages, and to put previously reached conclusions on trial.We may now provide answers to questions concerning the performance of different types of semantic representations (such as token embeddings vs. type embeddings, and topic models vs. vector space models), alignment methods and change measures.We provide a thorough analysis of the submitted results uncovering trends for models and opening perspectives for further improvements.In addition to this, the CodaLab website will remain open to allow any reader to directly and easily compare their results to the participating systems.We expect the long-term impact of the task to be significant, and hope to encourage the study of LSC in more languages than are currently studied, in particular less-resourced languages. SubtasksFor the proposed tasks we rely on the comparison of two time-specific corpora C 1 and C 2 .While this simplifies the LSC detection problem, it has two main advantages: (i) it reduces the number of time periods for which data has to be annotated, so we can annotate larger corpus samples and hence more reliably represent the sense distributions of target words; (ii) it reduces the task complexity, allowing * SH was affiliated with the