Contextual Embeddings for Ukrainian: A Large Language Model Approach to Word Sense Disambiguation

Yurii Laba, Volodymyr Mudryi, Dmytro Chaplynskyi, Mariana Romanyshyn, Oles Dobosevych · 2023

This research proposes a novel approach to the Word Sense Disambiguation (WSD) task in the Ukrainian language based on supervised finetuning of a pre-trained Large Language Model (LLM) on the dataset generated in an unsupervised way to obtain better contextual embeddings for words with multiple senses.The paper presents a method for generating a new dataset for WSD evaluation in the Ukrainian language based on the SUM dictionary.We developed a comprehensive framework that facilitates the generation of WSD evaluation datasets, enables the use of different prediction strategies, LLMs, and pooling strategies, and generates multiple performance reports.Our approach shows 77,9% accuracy for lexical meaning prediction for homonyms.

Read the paper · More papers on PaperTik