Non-Parametric Word Sense Disambiguation for Historical Languages

Enrique Manjavacas Arévalo, Lauren Fonteyn · 2022

Recent approaches to Word Sense Disambiguation (WSD) have profited from the enhanced contextualized word representations coming from contemporary Large Language Models (LLMs).This advancement is accompanied by a renewed interest in WSD applications in Humanities research, where the lack of suitable, specific WSD-annotated resources is a hurdle in developing ad-hoc WSD systems.Because they can exploit sentential context, LLMs are particularly suited for disambiguation tasks.Still, the application of LLMs is often limited to linear classifiers trained on top of the LLM architecture.In this paper, we follow recent developments in non-parametric learning and show how LLMs can be efficiently fine-tuned to achieve strong few-shot performance on WSD for historical languages (English and Dutch, date range: 1450-1950).We test our hypothesis using (i) a large, general evaluation set taken from large lexical databases, and (ii) a small real-world scenario involving an ad-hoc WSD task.Moreover, this paper marks the release of GysBERT, a LLM for historical Dutch.1 Both models can be accessed through the original huggingface repository through the following links https://huggingface.co/emanjavacas/ MacBERTh and https://huggingface.co/ emanjavacas/GysBERT.2 See Table 3 in the Appendix for an illustration of the structure of the lexical databases and an example of the sentences that are being classified.3 Note that for this approach to work, the true word sense

Read the paper · More papers on PaperTik