Token-level semantic typology without a massively parallel corpus
Barend Beekhuizen · 2025
This paper presents a computational method for token-level lexical semantic comparative research in an original text setting, as opposed to the more common massively parallel setting.Given a set of (non-massively parallel) bitexts, the method consists of leveraging pre-trained contextual vectors in a reference language to induce, for a token in one target language, the lexical items that all other target languages would have used, thus simulating a massively parallel set-up.The method is evaluated on its extraction and induction quality, and the use of the method for lexical semantic typological research is demonstrated.