SST-BERT at SemEval-2020 Task 1: Semantic Shift Tracing by Clustering in BERT-based Embedding Spaces
Vani Kanjirangat, Sandra Mitrović, Alessandro Antonucci, Fabio Rinaldi · 2020
Lexical semantic change detection (also known as semantic shift tracing) is a task of identifying words that have changed their meaning over time.Unsupervised semantic shift tracing, focal point of SemEval2020, is particularly challenging.Given the unsupervised setup, in this work, we propose to identify clusters among different occurrences of each target word, considering these as representatives of different word meanings.As such, disagreements in obtained clusters naturally allow to quantify the level of semantic shift per each target word in four target languages.To leverage this idea, clustering is performed on contextualized (BERT-based) embeddings of word occurrences.The obtained results show that our approach performs well both measured separately (per language) and overall, where we surpass all provided SemEval baselines. Problem SetupConsider two corpora C 1 and C 2 for a same language but associated with different time stamps (say, respectively, t 1 and t 2 > t 1 ).Let W be a set of target words occurring in both corpora.Each target word w ∈ W might assume multiple meanings, to be called senses, within the two corpora.A pool of experts annotated a representative amount of occurrences with their corresponding senses.The problem we consider is to characterize the semantic shift related to those senses from one corpus to the other without having access to the expert annotations.In particular, we address the two following two subtasks:• Subtask 1: Decide, for each w ∈ W, whether or not w gained or lost at least a sense between t 1 and t 2 .This is a binary decision task.We will denote this subtask as (S1).• Subtask 2: Define, for the elements of W, a measure of their degree of lexical semantic change between t 1 and t 2 and sort these elements consequently.This is a ranking task.This subtask will be referred to as (S2).We 1 describe two different methods able to address both subtasks.As both methods require a preprocessing step based on transformers, let us start from this preliminary operation.