Cross-Lingual Named Entity Recognition for the Ukrainian Language Based on Word Alignment
Artem A. Kramov, Sergiy D. Pogorilyy · 2022
Processing of low-resource languages is one of the most challenging tasks in the area of natural language processing. The lack of datasets makes it complex to create and train corresponding models that are available for high-resource languages. In this paper, the possibility of the usage of the word alignment mechanism for the performing of the cross-lingual tasks has been analyzed for the low-resource language by the example of the Ukrainian language for the solving of the named entity recognition task. Different state-of-the-art word alignment methods have been analyzed. The advisability of the usage of the pre-trained multilingual models for the word alignment of low-resource languages has been shown. The experimental verification of the effectiveness of the chosen word alignment method for the solving of the named entity recognition for a Ukrainian corpus has been performed with different options. The results obtained may indicate the possibility of the usage of the considered word alignment method for the solving of different sequence labeling tasks for the Ukrainian language by performing the cross-lingual mapping of the results for the target English language.