Found in Translation: Sourcing parallel corpora for low-resource language pairs

Hinrik Hafsteinsson, Steinþór Steingrímsson · Digital Humanities in the Nordic and Baltic Countries Publications · 2025

This paper describes the sourcing, processing, and application of parallel text data for Icelandic and Polish for the purpose of bilingual lexicon induction (BLI), demonstrating how a parallel corpus can be compiled for a low-to-medium resource language pair that has no available parallel data, by pivoting through a common language. We show the usefulness of the corpus by training and evaluating a machine translation (MT) model on the data. Iceland's linguistic landscape is evolving, with an increasing need for multilingual support due to the growing immigrant population. Polish, in particular, stands out as the language of the largest single minority in Iceland, underscoring the importance of this project.

Read the paper · More papers on PaperTik