Transformer-Based Multilingual Language Models in Cross-Lingual Plagiarism Detection

Tatevik Ter-Hovhannisyan, Karen Avetisyan · 2022

In this work, we study the effectiveness of 6 pretrained Transformer-based language models in the task of crosslingual sentence alignment. We evaluate and compare the models on 10 language pairs, defining the task as a binary classification of two input sentences in English and another Indo-European language. The main objective is to determine the best model in terms of processing speed, accuracy, and cross-lingual transferability. For the latter, we test the language models in three different settings: (i) when one fine-tuned model is used for all considered languages, (ii) when a separately fine-tuned model is used for each group of closely related languages, and (iii) when a separate model is fine-tuned for each language pair. For the experiments we use challenging translation identification datasets containing paraphrases and near-paraphrases.

Read the paper · More papers on PaperTik