Estonian Language Understanding: a Case Study on the COPA Task

Hele-Andra Kuulmets, Andre Tättar, Mark Fishel · Baltic Journal of Modern Computing · 2022

The lack of Estonian NLU datasets severely affects advancing Estonian-specific NLP research.With this paper we aim to relieve the issue by publishing a new Estonian NLU dataset EstCOPA.We benchmark the task on several Estonian and multilingual transformer based language models, including a novel Estonian-centric GPT (GPT4Est).Moreover, we evaluate different low-cost alternatives for creating training and test datasets and outline strategies for future Estonian language understanding research.

Read the paper · More papers on PaperTik